LLM評価
LLM評価ツールやベンチマーク、テスト基盤を探し、モデル品質の測定、出力比較、リスク検出、AIアプリの改善に役立てられます。
AIインフラ
このコレクションを探索製品

QAgent
AIテストアシスタント
QAgent is an AI agent testing and quality assurance platform for developers and agile teams. It connects to an agent through a webhook or endpoint, runs automated test cases, and evaluates responses for groundedness, policy adherence, prompt compliance, and related quality dimensions.

inferock-bench
AIゲートウェイとルーティング
inferock-bench is a local diagnostic proxy for tracking LLM API usage, provider-reported costs, failures, and billing-integrity signals. It helps developers inspect calls to OpenAI, Anthropic, Gemini Developer API, and pinned OpenRouter endpoints using locally stored receipts.





