Horizon is Labelbox’s RL environment product for post-training and evaluation. It helps AI teams build software-based environments, preference signals, and rubrics for long-horizon tasks in domains like reasoning, coding, computer use, and cybersecurity.

Horizon preview

Overview

Horizon is Labelbox’s RL environment product for post-training and evaluation. It focuses on generating software-based environments, preference signals, and rubrics for long-horizon AI tasks where reward signal quality matters.

The product is positioned for reasoning, tool use, and computer use across economically important knowledge-work domains. The site describes use cases such as autonomous AI research, scientific knowledge work, agent coding, multimodal work, voice with agentic tool use, computer use, and cybersecurity.

Core capabilities

Post-training environments

Build RL environments and preference-signal datasets for post-training and evaluation, with an emphasis on calibrated reward signals and pass@k targets.

Reasoning and agent-task support

Model tasks around reasoning, tool use, and computer use, including long-horizon workflows where verification and intermediate rewards matter.

Enterprise workflow simulation

Use WorldSim-style simulations to recreate enterprise environments such as GitLab, Jira, CRM, email, and chat with realistic business data.

Scenario generation at scale

Design configurable scenarios that generate diverse tasks at scale, including outages, workflow changes, and other long-tail situations.

Preference-signal labeling

Capture preference labels across agent trajectories to turn human judgment into structured comparisons for reward modeling.

Cross-domain coverage

Apply the same environment and evaluation approach across knowledge work domains such as autonomous research, software engineering, multimodal work, voice agents, computer use, and cybersecurity.

Common use cases

  • Enterprise workflow simulation

    Create simulations for enterprise knowledge work so agents can practice tasks in systems that resemble GitLab, Jira, CRM, email, and chat.

  • Reward modeling for long tasks

    Generate reward signals and rubrics for long-horizon tasks where models need intermediate verification rather than only final-answer scoring.

  • Agent coding and software engineering

    Build training and eval sets for software agents that debug, author PRs, and work through multi-step software engineering workflows.

  • Cybersecurity evaluation

    Construct scenarios for cybersecurity work, including attack and defense tasks with programmatic verification and adversarial edge cases.

  • Multimodal knowledge work

    Tune environments for knowledge work that spans text, images, charts, documents, and structured data within a single workflow.

Pros and Cons

Pros

  • Targets long-horizon tasks where reward signals and verification are difficult to design.
  • Supports multiple knowledge-work domains, including reasoning, coding, computer use, and cybersecurity.
  • Uses configurable scenarios and enterprise-like simulations to generate diverse task distributions.
  • Frames preferences and rubrics as structured signals for training and evaluation.

Cons

  • The source does not provide public pricing or packaging details for Horizon.
  • The site emphasizes environment and evaluation infrastructure, so teams would still need to integrate it into their own training and deployment workflow.

FAQ

What is Horizon used for?

Labelbox positions Horizon as RL environments and preference-signal infrastructure for post-training and evaluation. It is built for reasoning, tool use, computer use, autonomous research, scientific knowledge work, agent coding, and cybersecurity.

How does Horizon support training and evaluation?

The source describes Horizon as software-generated RL environments at scale, with calibrated reward signals and preference labels. It emphasizes scenarios, verification, difficulty progression, and credit assignment for long-horizon tasks.

Who is Horizon for?

Horizon is presented for frontier AI labs and teams working on long-horizon, economically important knowledge-work domains. The examples shown include scientific knowledge work, agent coding, computer use, and cybersecurity.

Does Horizon have public pricing?

The source does not list public pricing or packaging details for Horizon. It provides a contact path and a start-for-free call to action on the main site.

Quick Facts

Category
RL environments
Platform
Labelbox
Primary use
Post-training and evaluation
Site
labelbox.com
Related workflows
Reasoning, tool use, computer use