DEV Community

oooocean66 profile picture

oooocean66

Engineers working on GPU cloud infrastructure in Japan. We write hands-on notes on LLM fine-tuning, quantization, and inference benchmarks, plus real infra tradeoffs of running AI workloads at scale.

Joined Joined on 
Qwen3.8-27B: An Architecture and Hands-On Look at Alibaba's Compact Dense Model

Qwen3.8-27B: An Architecture and Hands-On Look at Alibaba's Compact Dense Model

Comments
24 min read

Want to connect with oooocean66?

Create an account to connect with oooocean66. You can also sign in below to proceed if you already have an account.

Already have an account? Sign in
Muse Glimmer 30B: An Architecture and Hands-On Look at Meta's New Mid-Range Model

Muse Glimmer 30B: An Architecture and Hands-On Look at Meta's New Mid-Range Model

Comments 1
17 min read
Trying "DFlash," a Diffusion-Model Approach to Parallel Draft-Token Generation, on Gemma

Trying "DFlash," a Diffusion-Model Approach to Parallel Draft-Token Generation, on Gemma

Comments 1
17 min read
MTP in Practice: Benchmarking Gemma's Speculative Decoding on a Real GPU

MTP in Practice: Benchmarking Gemma's Speculative Decoding on a Real GPU

Comments
13 min read
MTP Explained: How Qwen and Gemma Predict Several Tokens Ahead

MTP Explained: How Qwen and Gemma Predict Several Tokens Ahead

Comments 2
9 min read
loading...