DEV Community

Gemma

Gemma is a collection of lightweight, state-of-the-art open models built from the same technology that powers our Gemini models.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Gemma 4 on Amazon SageMaker: QAT Weights Decode 2.05x Faster Than bf16 on One L4

Gemma 4 on Amazon SageMaker: QAT Weights Decode 2.05x Faster Than bf16 on One L4

Comments
9 min read
Gemma 4 on an Amazon SageMaker Endpoint: AWS CLI, NVIDIA L4, and an MCP Server

Gemma 4 on an Amazon SageMaker Endpoint: AWS CLI, NVIDIA L4, and an MCP Server

Comments
10 min read
From smartphone to sovereign agent:How native on-device intelligence reshapes the personal computing device

From smartphone to sovereign agent:How native on-device intelligence reshapes the personal computing device

Comments
24 min read
Plain Gemma 4 26B vs Jev on One EC2 L4: 2.1 Points Behind Overall, Level on Yes/No, 4.5 Behind on Multiple Choice

Plain Gemma 4 26B vs Jev on One EC2 L4: 2.1 Points Behind Overall, Level on Yes/No, 4.5 Behind on Multiple Choice

1
Comments
23 min read
Google's QAT Gemma 4 26B-A4B on One TPU v6e: 15.6x the KV Cache and 1.9x the Throughput of FP8

Google's QAT Gemma 4 26B-A4B on One TPU v6e: 15.6x the KV Cache and 1.9x the Throughput of FP8

Comments 1
12 min read
Use Local LLM to Enable AI Enhancement for Speech-to-Text Dictation

Use Local LLM to Enable AI Enhancement for Speech-to-Text Dictation

Comments
5 min read
A 4 GB Laptop GPU vs a 6-Core CPU on Gemma 4, Re-Measured in ABBA Order: 4.1x

A 4 GB Laptop GPU vs a 6-Core CPU on Gemma 4, Re-Measured in ABBA Order: 4.1x

10
Comments 2
11 min read
Gemma 4 on Amazon SageMaker: QAT Weights Decode 2.05x Faster Than bf16 on One L4

Gemma 4 on Amazon SageMaker: QAT Weights Decode 2.05x Faster Than bf16 on One L4

7
Comments
9 min read
Gemma 4 on an Amazon SageMaker Endpoint: AWS CLI, NVIDIA L4, and an MCP Server

Gemma 4 on an Amazon SageMaker Endpoint: AWS CLI, NVIDIA L4, and an MCP Server

7
Comments
10 min read
Running a Jev-Style Decision Model on One TPU v6e: What Fits, What It Costs, and What Changes From a GPU

Running a Jev-Style Decision Model on One TPU v6e: What Fits, What It Costs, and What Changes From a GPU

10
Comments
13 min read
Plain Gemma 4 26B vs Jev on One EC2 L4: 2.1 Points Behind Overall, Level on Yes/No, 4.5 Behind on Multiple Choice

Plain Gemma 4 26B vs Jev on One EC2 L4: 2.1 Points Behind Overall, Level on Yes/No, 4.5 Behind on Multiple Choice

8
Comments
23 min read
Jev After Eight Days of Independent Tests: Level With Mid-Price LLMs, Behind the Frontier

Jev After Eight Days of Independent Tests: Level With Mid-Price LLMs, Behind the Frontier

9
Comments 1
17 min read
Gemma 4 on a Tesla T4, Part 2: The Minimum GCE VM and a Script to Drive It

Gemma 4 on a Tesla T4, Part 2: The Minimum GCE VM and a Script to Drive It

13
Comments
13 min read
Jev After Eight Days of Independent Tests: Level With Mid-Price LLMs, Behind the Frontier

Jev After Eight Days of Independent Tests: Level With Mid-Price LLMs, Behind the Frontier

3
Comments 2
17 min read
Why I Chose Gemma 4 E2B for Subra AI: The Reality of Running LLMs on Mid-Range Phones

Why I Chose Gemma 4 E2B for Subra AI: The Reality of Running LLMs on Mid-Range Phones

1
Comments
3 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.