Market Segment
Evaluation & Safety Vendors
Vendors evaluating model or agent capability, quality, and safety — including human and automated evaluation, verifier and reward-model development, and red teaming.
AI control and evaluation lab studying and deploying frontier agents through simulated benchmarks and real-world autonomous organizations.
AI evaluation company, formerly LMArena, operating a human-judgment evaluation platform and commercial evaluation services for model labs and enterprises.
Independent AI benchmarking company measuring full agents and models with realistic agentic task suites and harnesses.
Open-source environment lab building stateful agent benchmarks and evaluation infrastructure.
AI security company providing adversarial red-teaming and runtime protection for models and agents.
Independent evaluation platform building real-world, domain-specific benchmarks for AI models and agents.
About Evaluation & Safety
Evaluation & Safety vendors measure, test, and validate how models and agents actually behave, rather than providing the environments they act in or the infrastructure used to train them. Within RL.VENDORS, this includes benchmark and automated evaluation platforms, human-judgment evaluation, and adversarial red-teaming or security testing — work aimed at scoring, grading, or stress-testing a model or agent's behavior.
This differs from Environments & Simulations, which builds the task worlds agents operate in, and from Training Platforms, which focuses on the post-training workflow itself. Evaluation & Safety vendors instead sit downstream of both, assessing the resulting model or agent rather than producing it.
Buyers are typically model labs, AI product and safety teams, and enterprises evaluating an agent before deployment.