You've built and evaluated agents yourself, and you have opinions about what makes agent optimization actually work - evaluation harnesses, annotation flows, failure clustering, and the feedback loops that turn raw session data into product intelligence. Direct experience with the current agent-building and eval landscape (Cursor, Claude Code, Replit, LangSmith, Braintrust, Langfuse, agent frameworks, headless CMS) and strong opinions on what they get right and wrong.