DEV Community

Aleksei Romanov profile picture

Aleksei Romanov

Ex-Apple GenAI engineer, building an LLM post-training platform at g factor technologies

Joined Joined on  Personal website https://www.g-ftech.com
From RLHF to RLVR: The Evolution of Reward Signals and the Battle Against Reward Hacking

From RLHF to RLVR: The Evolution of Reward Signals and the Battle Against Reward Hacking

Comments 1
9 min read

Want to connect with Aleksei Romanov?

Create an account to connect with Aleksei Romanov. You can also sign in below to proceed if you already have an account.

Already have an account? Sign in
AsyncGRPO: Eliminating GPU Idle Bubbles in Environment-Heavy RL Post-Training

AsyncGRPO: Eliminating GPU Idle Bubbles in Environment-Heavy RL Post-Training

Comments 1
9 min read
Owning vs. Renting Intelligence: Why Enterprises Are Building Sovereign AI

Owning vs. Renting Intelligence: Why Enterprises Are Building Sovereign AI

Comments 2
9 min read
LoRA & DoRA: The Math, Memory, and Trade-offs

LoRA & DoRA: The Math, Memory, and Trade-offs

2
Comments 3
8 min read
Multi-Reward Reinforcement Learning for LLM Agents: Comparing PPO, GRPO, DAPO, and GDPO

Multi-Reward Reinforcement Learning for LLM Agents: Comparing PPO, GRPO, DAPO, and GDPO

Comments 2
8 min read
SFT vs. RL: What Changes Inside the Model?

SFT vs. RL: What Changes Inside the Model?

Comments 6
6 min read
Latent-GRPO: Reinforcement Learning in Continuous Thought Space

Latent-GRPO: Reinforcement Learning in Continuous Thought Space

Comments 2
9 min read
High-Throughput LLM Inference & Training: A Deep Dive into vLLM

High-Throughput LLM Inference & Training: A Deep Dive into vLLM

2
Comments 1
9 min read
loading...