RL.VENDORS

Market Segment

Training Platforms Vendors

Vendors providing platforms that orchestrate the RL or post-training process itself — managed training APIs, algorithm support, environment and reward integration, and compute orchestration at scale.

Focus Area
Domain
Service / Capability
Company Type
Sort by
AgileRL logo
AgileRLTraining Platforms

AgileRL provides enterprise RL training and RLOps for building, tuning, deploying, and continually improving specialized AI agents at scale.

RL EnvironmentsRLHF / Post-trainingAgent Infrastructure
Applied Compute logo
Applied ComputeTraining Platforms

Builds enterprise-specific intelligence using targeted reinforcement learning over integrated environments plus deployment infrastructure.

Enterprise AgentsRL EnvironmentsRLHF / Post-training
Ashr logo
AshrTraining Platforms

Ashr provides enterprise post-training and continual learning for open-weight models using production feedback, fine-tuning, evaluation, and serving.

Enterprise AgentsRLHF / Post-trainingEvaluations
Castform logo
CastformTraining Platforms

Castform provides managed RL post-training for open-weight models with data preparation, rewards, tools, GPU training, evaluation, and deployment.

Enterprise AgentsRL EnvironmentsRLHF / Post-training
Osmosis logo
OsmosisTraining Platforms

Forward-deployed reinforcement-learning platform for task-specific agent/model post-training with custom tools, reward functions, rollouts, and retraining loops.

RLHF / Post-trainingAgent InfrastructureSoftware Engineering
Oumi logo
OumiTraining Platforms

Open-source and hosted platform for building, evaluating, post-training, deploying, and continuously improving specialized AI models.

Training DataRLHF / Post-trainingSoftware Engineering
Thinking Machines Lab / Tinker logo
Thinking Machines Lab / TinkerTraining Platforms

Model-training API and cookbook with composable RL environments, GRPO pipelines, tool-use training, and multi-agent RL recipes.

RLHF / Post-trainingAgent InfrastructureSoftware Engineering
TrainLoop logo
TrainLoopTraining Platforms

TrainLoop trains reasoning models and agents for long-horizon enterprise tasks using reinforcement learning, evaluation, and continual learning.

Enterprise AgentsRLHF / Post-trainingEvaluations
ZE
Zero Proof LabsTraining Platforms

Zero Proof Labs provides agent simulations, evaluation, training data, and hosted SFT, GRPO, DPO, and reward-model post-training workflows.

RL EnvironmentsTraining DataRLHF / Post-training

About Training Platforms

Training Platforms provide the infrastructure and workflows used to actually train or post-train models and agents, rather than only supplying the raw compute or environments underneath that process. A platform in this category may combine several capabilities — supervised fine-tuning, reinforcement learning, reward-function or verifier design, evaluation loops, and managed training infrastructure — into a single offering aimed at improving model or agent behavior end to end.

This differs from RL Infrastructure, which centers on the underlying compute, execution, orchestration, and sandboxing that training and evaluation runs depend on, rather than the training workflow itself. The distinction is one of emphasis, not a strict boundary, since some vendors span both.

RL.VENDORS assigns each vendor exactly one primary market segment, based on its main commercial positioning, even when its public offering spans several of these categories.