Agent Evaluation Studio
A small laboratory for asking a useful question: how do we know when an AI agent is actually doing a good job?
PythonLLM evaluationAgentsStructured data
Read the story I like projects that begin with genuine curiosity and end somewhere useful. Here are a few explorations across agents, generative AI, evaluation and data.
A small laboratory for asking a useful question: how do we know when an AI agent is actually doing a good job?
A research companion that gathers, compares and distills sources without pretending uncertainty has disappeared.
A visual data essay about the ordinary distances between Brazil and Italy: weather, words, time and daily rituals.