OutEvalAI LogoOutEvalAI

OutEvalAI Documentation

Evaluate AI agents with personas, scenarios, and automated reports using OutEvalAI.

OutEvalAI is a workspace for evaluating AI agents before they go live. Product and engineering teams configure usecases under test, define personas and scenarios, and review automated pass/fail reports with transcripts and improvement suggestions.

Use these guides to move from a first evaluation run to a repeatable quality workflow for agent changes.

Choose where to begin

Start with the evaluation model, configure an agent under test, then integrate through the API or SDKs.

Product guides

Focused documentation for the evaluation product surface.

Recommended workflow

A practical loop for shipping safer agent changes.

  1. 1

    Define the agent under test

    Create or select a usecase that represents the conversation you want to evaluate.

  2. 2

    Create personas and scenarios

    Describe who the simulated caller is and what success looks like for each test case.

  3. 3

    Run AI-to-AI evaluation

    Execute live tests against your usecase and capture transcripts plus structured outcomes.

  4. 4

    Review pass/fail reports

    Inspect automated scoring, failure reasons, and suggestions before shipping changes.

  5. 5

    Iterate and re-test

    Refine prompts and agent configuration, then re-run scenarios to confirm improvements.

How is this guide?