Agent Evaluation Studio
A small laboratory for asking a useful question: how do we know when an AI agent is actually doing a good job?
Why build this?
Agent demos can look convincing while hiding fragile reasoning, inconsistent tool use and silent failures.
What I made
A repeatable evaluation workflow that traces decisions, tests realistic scenarios and turns qualitative behavior into evidence a team can discuss.