Evals and traces around a feature that is already live.
Right now you cannot tell the difference between a bad day and a regression. We build the eval set that defines what good means for your flow, and the tracing that lets you explain a specific bad answer after a customer reports it.