Experiment tracking
Track machine learning experiments and model runs, then review them in the W&B platform for comparison and analysis.
Machine learning and AI platform for experiment tracking, artifacts, GenAI tracing, and evaluations.
Weights & Biases is a developer tool for machine learning and AI application teams. It provides a platform for tracking experiments, managing model and dataset artifacts, and observing GenAI systems across the development and production workflow.
The product spans multiple deployment styles, including hosted SaaS, dedicated cloud, and customer-managed options. The pricing pages also show separate paths for personal use, teams, and enterprise customers, with paid support and security features available on higher tiers.
Track machine learning experiments and model runs, then review them in the W&B platform for comparison and analysis.
Version datasets and models, manage artifacts, and follow lineage across the model lifecycle.
Build collaborative dashboards and reports to share results and progress with teammates.
Log GenAI traces, inputs, outputs, and metadata for evaluation and production monitoring during development.
Run evaluations, compare recipes side by side, and inspect accuracy, latency, and token usage for LLM applications.
Use access controls, team-based permissions, service accounts, audit logs, and SSO on supported plans.
Track runs, compare results, and keep artifacts organized while iterating on models and experiments.
Record traces, evaluations, and production monitoring signals for LLM applications and agent workflows.
Share dashboards, reports, and artifacts so researchers and engineers can review results together.
Deploy in a SaaS, dedicated, or customer-managed environment when security, isolation, or data residency requirements matter.
Use the personal, academic, or trial paths to evaluate the platform before committing to a broader rollout.
Weights & Biases bills the Pro plan monthly or annually upfront, depending on the selection. Model seats are prorated if added mid-term, while Weave data ingestion, W&B Inference, and storage are billed monthly in arrears based on usage. Enterprise plans are invoiced annually upfront.
The pricing page says the free plan is designed for personal development of AI applications and models, and also offers personal and academic options for research and local use. Corporate use is not allowed on the Personal plan.
Enterprise plans add single-tenant or customer-managed deployment options, security controls such as SSO, automated user provisioning, audit logs, and customer-managed encryption key support, depending on the deployment option.
The support page lists Standard, Standard Plus, and Premium support packages. Standard Plus adds onboarding, a dedicated success team, and Slack or Microsoft Teams support, while Premium adds faster response times, 24/7 coverage, and collaboration features such as feature prioritization and roadmap sessions.
The deployment options page describes SaaS Cloud, Dedicated Cloud, and Customer-Managed deployments. W&B also documents locally hosted server options on the pricing page for personal use and an enterprise trial for running a server on your own infrastructure.
Traffic data is for reference only.
| May | 2475536 |
|---|---|
| Jun | 2045155 |
| Jul | 2083772 |
Traffic analytics are not available yet.
Traffic analytics are not available yet.
www.lyzr.ai
OpenController is Lyzr’s control plane for discovering, evaluating, governing, and monitoring AI agents, models, tools, data, and workflows across an enterprise AI estate. It is intended for teams managing agents across clouds, frameworks, runtimes, and environments.
www.langchain.com
LangSmith is an observability and evaluation platform for AI agents and LLM applications. It helps development and production teams trace agent behavior, monitor quality and cost, investigate failures, and evaluate changes.
www.galileo.ai
Galileo is an AI observability and evaluation platform for testing, debugging, and governing LLM and agent systems across development and production. It helps teams turn evaluation results into production guardrails and monitor AI behavior at scale.
www.tensorzero.com
TensorZero is an open-source LLMOps platform for building and operating production-grade LLM applications. Its stated scope combines an LLM gateway with observability, evaluation, optimization, and experimentation tools.
arize.com
Phoenix is an open-source, local-first platform for tracing, evaluating, experimenting with, and improving AI applications and agents. It helps AI engineers inspect agent behavior, assess output quality, and test changes before deployment.
wandb.ai
W&B Weave is an observability and evaluation platform for production AI agents and applications. It helps teams trace agent behavior, evaluate changes, inspect prompts and models, and monitor production interactions.