Evals · Check #1 — sensor fidelity
Everything the sensors captured this week, joined with the threads each entity carries. Grade each item — and report anything that should have been captured but wasn’t. Verdicts persist and tune the system.
Something missing?
Check #2 — thread enrichment
Per open thread: the context the system attached (notes / meetings / mail) and any question it proactively asked. Is the grounding right?