Reinforcement learning environments
Construct durable environments for measuring agentic capability in work settings with long horizons, natural instructions, realistic tools, and domain-specific judgment.
Datacurve builds data and evaluation infrastructure for frontier AI, including datasets, benchmarks, environments, and expert trajectories for long-horizon reasoning.
Datacurve describes itself as the data engine for frontier AI and as infrastructure for research and data collection that helps teach models to handle long-horizon work. Its site focuses on the kinds of data needed for reasoning, software engineering, and data science, especially where judgment, iteration, and partial progress matter.
The product page presents Datacurve as a provider of datasets, benchmarks, evaluations, reinforcement learning environments, and expert trajectories. The public research page adds Deep SWE, a benchmark for long-horizon software engineering, which shows the company applying that approach to evaluating and measuring agent performance on more realistic engineering tasks.
Construct durable environments for measuring agentic capability in work settings with long horizons, natural instructions, realistic tools, and domain-specific judgment.
Collect tasks that stretch over hours or days and preserve ambiguity, partial progress, and recovery so model training reflects real work instead of simplified prompts.
Provide prebuilt datasets that are curated for signal, reviewed for quality, and structured to fit directly into a training stack.
Build benchmarks that aim to capture task-faithful, domain-sensitive lift rather than only optimizing a single numeric score.
Capture full expert execution traces, including tool calls, checks, pivots, and recoveries, so agents can learn process as well as outcome.
Create demonstrations for supervised fine-tuning using bespoke tooling that lets experts work naturally while preserving the operating judgment behind the work.
Train or evaluate coding agents on longer engineering jobs where a correct outcome depends on many intermediate steps, tool uses, and recoveries.
Assemble datasets and evaluations for model development in scientific or analytical settings where task structure and judgment matter.
Create realistic reinforcement learning environments that reflect the complexity of real systems, including partial progress and domain-specific decisions.
Use expert execution traces and demonstrations to teach models how experts work through tasks, not just what final answers look like.
Datacurve positions itself as infrastructure for research and data collection for frontier AI. Its products page focuses on reinforcement learning environments, long-horizon tasks, OTS datasets, benchmarks, agent trajectories, and supervised fine-tuning data.
The site highlights software engineering, data science, cyber security, machine learning, and research as domains it works in, with a stated emphasis on long-horizon reasoning and verifiable work.
The research page introduces Deep SWE, a benchmark for long-horizon software engineering, and says Datacurve publishes benchmarks, datasets, and evaluations from its own research.
The pricing page in the collected evidence returns a 404, so the site does not currently expose a stable pricing page in the provided sources.