Building the RL ecosystem
Find the companies
powering real-world RL.
A curated directory of companies building environments, data, evaluations, infrastructure, and related capabilities for reinforcement learning and agent systems.
Applied research lab building training data, evaluation datasets, and reinforcement-learning environments around expert professional workflows.
AgileRL provides enterprise RL training and RLOps for building, tuning, deploying, and continually improving specialized AI agents at scale.
AI control and evaluation lab studying and deploying frontier agents through simulated benchmarks and real-world autonomous organizations.
RL data lab programmatically generating environments, tasks, and verifiers from real-world data for post-training and evaluation.
Post-training and evaluation company combining subject-matter experts with data workflows for LLM evaluation, RLHF, supervised fine-tuning, and scalable oversight.
Ray-based distributed AI platform supporting large-scale LLM reinforcement learning, multi-turn agent RL, tool use, reward computation, and post-training.
AI data provider offering end-to-end reinforcement-learning environment design, human data, evaluation, and verifier/reward construction for agentic AI.
Builds enterprise-specific intelligence using targeted reinforcement learning over integrated environments plus deployment infrastructure.
AI evaluation company, formerly LMArena, operating a human-judgment evaluation platform and commercial evaluation services for model labs and enterprises.
Builds production-fidelity RLVR environments and long-horizon benchmarks in cyber, SRE, compilation, and STEM.
Independent AI benchmarking company measuring full agents and models with realistic agentic task suites and harnesses.
Ashr provides enterprise post-training and continual learning for open-weight models using production feedback, fine-tuning, evaluation, and serving.
Builds long-horizon RL environments and post-training datasets for coding, computer use, knowledge work, and research.
Open-source environment lab building stateful agent benchmarks and evaluation infrastructure.
Applied AI research lab building reinforcement-learning environments and infrastructure, evaluation, and data-curation tooling for AI agents.
Castform provides managed RL post-training for open-weight models with data preparation, rewards, tools, GPU training, evaluation, and deployment.
Builds high-fidelity RL environments, evaluations, and human computer-use trajectory datasets for frontier AI research and post-training.
Builds realistic simulation labs, RL environments, verifiers, benchmarks, and training data for frontier agents.
Provides specialized AI cloud infrastructure, including managed and serverless reinforcement-learning infrastructure for training AI agents.
Open-source infrastructure provider developing cross-operating-system desktop drivers, virtualized sandboxes, and computer-use benchmark suites.
Builds training and evaluation data, including reinforcement-learning environments and agent trajectories, for frontier model development.
Provides isolated full-computer sandboxes for AI agents across containers, VMs, Windows, GPU, and macOS environments.
Builds financial process data, graders, sealed evaluations, and long-horizon RL environments from real institutional-finance decisions.
Digital-twin simulation vendor whose Falcon platform supports synthetic data generation, robotics testing, autonomy validation, and reinforcement-learning workflows.
Cloud sandbox infrastructure providing isolated, code-executable environments for AI agents.
Builds multi-step RL environments from financial market data for evaluation and post-training of research agents.
Builds RL environments that simulate production engineering systems and software-lifecycle complexity.
Applied research lab building simulated environments and real-world scenarios for training and evaluating agents.
Research company building open reward and RL-environment infrastructure, including OpenReward.
Turns games into interactive learning environments with verifiable signals for training and evaluating frontier models.
AI security company providing adversarial red-teaming and runtime protection for models and agents.
Builds computer- and tool-use RL environments, datasets, evaluations, and verifiers for economically valuable financial-services workflows.
Handshake's frontier-AI business spanning expert human data, evaluations, verifier research, and RL-environment work.
hiloop provides automated-research infrastructure with forkable compute, sandboxes, experiment lineage, observability, and training orchestration.
Platform for building, running, and evaluating reinforcement-learning environments and agent tasks.
Builds reinforcement-learning environments, computer-use models, and evaluation tooling for enterprise AI deployment.
EXL-owned AI training and evaluation business providing expert-led model training, evaluation, RL-related data, and multimodal data services.
Builds high-fidelity, difficult, verifiable RL environments for autonomous cybersecurity capabilities.
Provides RL environments, expert data, annotation, and evaluation infrastructure for model builders and enterprises across multiple professional verticals.
Provides reinforcement-learning data, environments, evaluations, and expert-feedback infrastructure for frontier AI teams.
Software company building reinforcement learning environments and software engineering evaluations for frontier coding agents.
Expert marketplace providing human datasets, benchmarks, and reinforcement-learning environments for professional-workflow AI training.
Builds RL environments, agent evaluations, datasets, and long-horizon enterprise benchmarks.
Data lab combining expert human data, Realm RL environments, contextual agent evaluations, and robotics data.
Provides high-concurrency sandboxes for AI code execution and RL rollouts, including coding-agent training environments.
Cloud infrastructure company providing programmable persistent VM devboxes and branching environments for coding agents, computer-use agents, development, and testing.
Managed reinforcement-learning evaluation environments for testing tool-using agents across reproducible task suites.
Turns real enterprise data into anonymized digital twins and expert-level RL environments for agents.
Forward-deployed reinforcement-learning platform for task-specific agent/model post-training with custom tools, reward functions, rollouts, and retraining loops.
Open-source and hosted platform for building, evaluating, post-training, deploying, and continuously improving specialized AI models.
Builds simulation and evaluation infrastructure for training and testing AI agents, including Digital World Models, generative RL environments, benchmarks, and agentic supervision tools.
Builds long-horizon coding RL environments from licensed private production repositories for frontier-model post-training.
AI research-engineering company building high-quality RL training environments and reward functions for real-world machine-learning research tasks.
Open infrastructure stack for reinforcement-learning environments, hosted training, evaluation, and compute.
Human-data and evaluation platform providing verified participants and domain experts for AI training, alignment, safety testing, and model evaluation.
Expert-data and evaluation company turning proprietary domain data into verified RL training sets, sandboxed environments, and adaptive evaluations.
Builds RL environments, evaluations, and verifiable training data for coding and computer-use agents, including software-world simulations, Terminal-Bench-style tasks, and long-horizon tool gyms.
Provides expert human data, RLHF, evaluations, agent training trajectories, and custom RL environments for frontier and domain-specific AI systems.
Physical-AI data infrastructure company, formerly Voxelmaps, providing world and human datasets for autonomous vehicles, humanoid robots, drones, and intelligent machines.
Provides VM-based Devbox sandboxes where AI agents can safely execute code, use files, APIs, and browsers.
AI data company providing human-in-the-loop training data, model evaluation, annotation, and preference/alignment workflows for advanced AI systems.
Data infrastructure company offering training data, evaluation platforms, and reinforcement learning environments for AI model development.
Develops data, environments, and evaluation infrastructure, including expert-data workflows, for specialized and agentic AI systems.
Human-data and RL-environment company providing expert data, environment tooling, trajectories, verifiers, evaluations, and post-training workflows for frontier AI teams.
Human-data and evaluation provider offering expert data, RLHF, and reinforcement-learning environments for frontier AI systems.
Model-training API and cookbook with composable RL environments, GRPO pipelines, tool-use training, and multi-agent RL recipes.
AI data company building RL gyms, virtual environments, trajectories, human-feedback data, evaluations, and safety datasets for AI agents.
TrainLoop trains reasoning models and agents for long-horizon enterprise tasks using reinforcement learning, evaluation, and continual learning.
Provides datasets, reinforcement-learning environments, expert contributors, and benchmarks for frontier AI model training and evaluation.
Provides reasoning datasets, benchmarks, and verifier-backed RL gyms for mathematics, algorithms, science, and code.
Independent evaluation platform building real-world, domain-specific benchmarks for AI models and agents.
Simulation infrastructure that recreates production users, systems, APIs, and data for agent evaluation, benchmarking, RL, and SFT.
Reinforcement-learning company building self-play task generation and environment infrastructure that turns proprietary data and evaluations into new training environments.
Zero Proof Labs provides agent simulations, evaluation, training data, and hosted SFT, GRPO, DPO, and reward-model post-training workflows.
Our approach
Public sources. Unknowns stay unknown.
We use a consistent, source-driven research process, keep unverifiable information explicit, and do not use rankings or pay-to-rank placement.