Benjamin rlhf
Vedika API · Pune, India
Job Description
Hiring: Benjamin RL
We're training the next generation of Vedika models.
This role sits directly inside RL and post-training research: designing how the model learns after pretraining, how it improves from interaction, how it handles long-horizon tasks, and how we push capability beyond standard instruction tuning.
You'll work on:
RL training for next-generation Vedika models GRPO, PPO, DPO and newer post-training methods Reward models, process rewards and verifiers Long-horizon reasoning and agent trajectories Tool-use and computer-use reinforcement Self-improvement and synthetic training loops Multi-turn behaviour and memory training Failure mining from model trajectories Evaluation systems for reasoning, autonomy and reliability Research experiments that can become part of the next model generationCompensation:
₹2.6 LPA fixed
₹3.6 LPA CTC
Work mode: Fully remote
You'll get:
Mac for development Claude Codex Serious compute and research infrastructure ₹10 L–₹50 L+ yearly AI/token spend available across the team and experimentsThis is not a role for someone whose idea of model work ends at prompting or basic fine-tuning.
We want someone who can understand a training run, break it, diagnose it, redesign it and make the next model measurably better.
Strong Py Torch, RL fundamentals, post-training, distributed training and hands-on experimentation matter far more than credentials.
Role: Benjamin RL
Vedika — Next Generation Models
Send your work, experiments, papers, repos or anything you trained that genuinely got better.
Details
| Company | Vedika API |
| Location | Pune, India |
| Type | FULL TIME |
| Niche | general |
| Experience | permanent |
