← All projects
01AI Evaluation2026

Agent Evaluation Studio

A small laboratory for asking a useful question: how do we know when an AI agent is actually doing a good job?

The question

Why build this?

Agent demos can look convincing while hiding fragile reasoning, inconsistent tool use and silent failures.

The work

What I made

A repeatable evaluation workflow that traces decisions, tests realistic scenarios and turns qualitative behavior into evidence a team can discuss.

View on GitHub
Continue exploringMore things I’ve built