Weekly traces hour for agent quality
Review real AI traces in a standing weekly session and turn the sharpest failures and good catches into eval cases.
Why this can grow a startup
Agent failures are often too specific and situational to notice through synthetic tests alone. A recurring trace review forces the team to watch what users actually asked, where the model drifted, and which interventions felt helpful. Converting those observations into eval cases compounds the learning instead of letting each debugging session disappear into chat history.
Company example
PostHog says the team runs a weekly traces hour, manually reviews real sessions with ratings, and then turns both bad failures and strong interventions into evals so future model or prompt changes do not regress those behaviors.
Source and metric
Source: PostHog Newsletter · Browse PostHog Newsletter tactics
PostHog reviews real rated agent sessions weekly and uses those findings to create future eval cases.
Source discovered: May 26, 2026
When to use it
Use this when Product, Retention, Support is relevant to ai products, retention, quality and you can run a bounded test with a low budget.
When not to use it
Do not use it as a substitute for customer evidence, a clear owner, or a measurable stop condition. Local platform rules and market behavior still need checking.
Founder checklist
- Read PostHog Newsletter and identify what is directly supported.
- Choose one channel context: Product, Retention, Support.
- Define the test around PostHog reviews real rated agent sessions weekly and uses those findings to create future eval cases..
- Set an owner, evidence window, and stop condition before launch.
Explore the context
Apply this with an operator
Connect activation, customer value, retention, and referral into one measurable loop.