LLM Evaluation
Explore LLM evaluation tools, benchmarks, and testing platforms to measure model quality, compare outputs, detect risks, and improve AI applications.
AI Infrastructure
Explore this collectionProducts

QAgent
AI Testing Assistant
QAgent is an AI agent testing and quality assurance platform for developers and agile teams. It connects to an agent through a webhook or endpoint, runs automated test cases, and evaluates responses for groundedness, policy adherence, prompt compliance, and related quality dimensions.

inferock-bench
AI Gateway And Routing
inferock-bench is a local diagnostic proxy for tracking LLM API usage, provider-reported costs, failures, and billing-integrity signals. It helps developers inspect calls to OpenAI, Anthropic, Gemini Developer API, and pinned OpenRouter endpoints using locally stored receipts.





