Post-training environments
Build RL environments and preference-signal datasets for post-training and evaluation, with an emphasis on calibrated reward signals and pass@k targets.
Horizon is Labelbox’s RL environment product for post-training and evaluation. It focuses on generating software-based environments, preference signals, and rubrics for long-horizon AI tasks where reward signal quality matters.
The product is positioned for reasoning, tool use, and computer use across economically important knowledge-work domains. The site describes use cases such as autonomous AI research, scientific knowledge work, agent coding, multimodal work, voice with agentic tool use, computer use, and cybersecurity.
Build RL environments and preference-signal datasets for post-training and evaluation, with an emphasis on calibrated reward signals and pass@k targets.
Model tasks around reasoning, tool use, and computer use, including long-horizon workflows where verification and intermediate rewards matter.
Use WorldSim-style simulations to recreate enterprise environments such as GitLab, Jira, CRM, email, and chat with realistic business data.
Design configurable scenarios that generate diverse tasks at scale, including outages, workflow changes, and other long-tail situations.
Capture preference labels across agent trajectories to turn human judgment into structured comparisons for reward modeling.
Apply the same environment and evaluation approach across knowledge work domains such as autonomous research, software engineering, multimodal work, voice agents, computer use, and cybersecurity.
Create simulations for enterprise knowledge work so agents can practice tasks in systems that resemble GitLab, Jira, CRM, email, and chat.
Generate reward signals and rubrics for long-horizon tasks where models need intermediate verification rather than only final-answer scoring.
Build training and eval sets for software agents that debug, author PRs, and work through multi-step software engineering workflows.
Construct scenarios for cybersecurity work, including attack and defense tasks with programmatic verification and adversarial edge cases.
Tune environments for knowledge work that spans text, images, charts, documents, and structured data within a single workflow.
Labelbox positions Horizon as RL environments and preference-signal infrastructure for post-training and evaluation. It is built for reasoning, tool use, computer use, autonomous research, scientific knowledge work, agent coding, and cybersecurity.
The source describes Horizon as software-generated RL environments at scale, with calibrated reward signals and preference labels. It emphasizes scenarios, verification, difficulty progression, and credit assignment for long-horizon tasks.
Horizon is presented for frontier AI labs and teams working on long-horizon, economically important knowledge-work domains. The examples shown include scientific knowledge work, agent coding, computer use, and cybersecurity.
The source does not list public pricing or packaging details for Horizon. It provides a contact path and a start-for-free call to action on the main site.
トラフィックデータは参考情報としてご利用ください。
opentrain.ai
OpenTrain AIは、事前審査済みのAIトレーナー、データラベラー、ドメイン専門家を採用できるマーケットプレイス兼マネージドサービスです。RLHF、評価、レッドチーム、アノテーション、エージェントワークフローに対応します。
www.lyzr.ai
OpenController is Lyzr’s control plane for discovering, evaluating, governing, and monitoring AI agents, models, tools, data, and workflows across an enterprise AI estate. It is intended for teams managing agents across clouds, frameworks, runtimes, and environments.
computearena.ai
ComputeArena is a community benchmark directory for measuring local AI model throughput across chips, quantisations, and runtimes. It helps developers run offline benchmarks, inspect signed reports, and compare decode and prefill performance on compatible workloads.
context.ai
Context は、顧客インフラ上で AI エージェントを構築・デプロイ・改善できる企業向けプラットフォームです。ワークスペース、ランタイム、コンテキスト、評価ツール、コネクタ、IdP ベースのアクセス制御を備えています。
www.langchain.com
LangSmith is an observability and evaluation platform for AI agents and LLM applications. It helps development and production teams trace agent behavior, monitor quality and cost, investigate failures, and evaluate changes.
deepeval.com
DeepEval is an open-source LLM evaluation framework for testing and benchmarking AI applications. It helps developers run pytest-native evaluations, score outputs and agent traces, and iterate on systems across text, image, audio, and voice workflows.