The instruction layer measures itself.
Every capability lives or dies by whether agents are actually told to use it — so Firekeep scores its own instructions against what sessions actually did. The Autopilot tab's compliance table shows, per instruction, how many sessions recalled before answering, recorded as they went, declared consequential actions — computed deterministically from the replay record, and honest about what it can claim: this measures behavior, not whether the behavior helped. The first live read, on the maintainer's own deployment (32 sessions, window split in halves by time): recall-before-answering went 44% → 69%, record-as-you-go 31% → 63%, recalled-knowledge-visibly-used 13% → 44%. Behavior moving, measured — with the quality question deliberately left to the A/B rounds ahead: instruction rewrites drafted by your own agents, approved by you, validated across real sessions. And because a rate is only as honest as its denominator, sessions now carry receipts of which instruction text actually reached them — content-hashed per runtime, with everything unverifiable reported as unknown rather than counted against anyone.
per-instruction compliance from replay (shipped) · trend over time (shipped) · exposure receipts + per-runtime slicing (shipped) · fleet-drafted rewrites under human verdict · A/B validation — roadmap