DEV Community

#vllm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Gemma 4 on Amazon SageMaker: The NVIDIA T4 Decodes at 0.8x of the L4 With the Same Answers

Gemma 4 on Amazon SageMaker: The NVIDIA T4 Decodes at 0.8x of the L4 With the Same Answers

15
Picked as gem Comments 1
9 min read
Gemma 4 on Amazon SageMaker: 4-Bit Embeddings Decode up to 1.39x Faster on One L4

Gemma 4 on Amazon SageMaker: 4-Bit Embeddings Decode up to 1.39x Faster on One L4

1
Comments
11 min read
Gemma 4 on Amazon SageMaker or a VM? The Same Model Server at 1.40x the Price

Gemma 4 on Amazon SageMaker or a VM? The Same Model Server at 1.40x the Price

1
Comments
8 min read
Gemma 4 on a Tesla T4, Part 3: Int4 Embeddings Serve E2B in 2.86 GiB at 2.30x bf16

Gemma 4 on a Tesla T4, Part 3: Int4 Embeddings Serve E2B in 2.86 GiB at 2.30x bf16

9
Comments
11 min read
Gemma 4 on Amazon SageMaker: QAT Weights Decode 2.05x Faster Than bf16 on One L4

Gemma 4 on Amazon SageMaker: QAT Weights Decode 2.05x Faster Than bf16 on One L4

Comments
9 min read
Gemma 4 on Amazon SageMaker: The NVIDIA T4 Decodes at 0.8x of the L4 With the Same Answers

Gemma 4 on Amazon SageMaker: The NVIDIA T4 Decodes at 0.8x of the L4 With the Same Answers

1
Comments
9 min read
Run vLLM on Kubernetes with Minikube, WSL2 and NVIDIA GPU

Run vLLM on Kubernetes with Minikube, WSL2 and NVIDIA GPU

Comments
13 min read
KV Cache on 16 GB GPUs: Making Long Context Actually Fit

KV Cache on 16 GB GPUs: Making Long Context Actually Fit

Comments 1
22 min read
Inside vLLM: Following One Request from the API to GPU Execution

Inside vLLM: Following One Request from the API to GPU Execution

1
Comments 2
24 min read
Gemma 4 on Amazon SageMaker: QAT Weights Decode 2.05x Faster Than bf16 on One L4

Gemma 4 on Amazon SageMaker: QAT Weights Decode 2.05x Faster Than bf16 on One L4

7
Comments
9 min read
Gemma 4 on a Tesla T4, Part 2: The Minimum GCE VM and a Script to Drive It

Gemma 4 on a Tesla T4, Part 2: The Minimum GCE VM and a Script to Drive It

13
Comments
13 min read
vLLM v0.28.0: the breaking change small GPU users must read

vLLM v0.28.0: the breaking change small GPU users must read

Comments
5 min read
Serving Gemma 4 on an AMD MI300X: What $1.99 an Hour Buys

Real-world benchmarks and MCP tool management tips

Serving Gemma 4 on an AMD MI300X: What $1.99 an Hour Buys

14
Comments 7
9 min read
Can vLLM Run GGUF? Yes — on GPU Only

Can vLLM Run GGUF? Yes — on GPU Only

1
Comments 1
4 min read
Gemma 4 on a Tesla T4: QAT Weights Decode 1.79x Faster Than bf16

Gemma 4 on a Tesla T4: QAT Weights Decode 1.79x Faster Than bf16

8
Comments 1
9 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.