4.2 Improve and survive.
Builds: Operating review and recovery plan with 30, 90, 180, 365 dates; Reinforcement runway; Executive status report.
Source: Interview 10. The interview numbers are the playbook's reference labels; this page is the section.
Personal
A second tool runs your week from the exported vault and nothing is lost.
Professional
A colleague resumes a recurring process from the records during your absence.
Artifact · Operating review, maintenance plan, recovery plan, and continuity handoff.
Why it exists
A system becomes durable when it can be corrected, backed up, restored, handed over, and resumed without reconstructing context from old chats. This interview also owns the calendar: the reinforcement runway touchpoints and the 30/90/180/365 reviews are scheduled here, with real dates, before the interview closes.
Questions
- Who owns the operating system?
- Which workflow is being reviewed first?
- What is the operating cadence?
- When was it last used in real work?
- Which three to five measures can change decisions? (Standard set: minutes per run, review minutes per run, error/rework rate, accepted-output rate, stale-source errors.)
- What is the baseline? (Pull from the intake brief; pull current numbers from the Evidence Log.)
- What has improved, in numbers?
- What failures or rework occurred?
- What caused each failure?
- Which maintained record, rule, or process was corrected?
- Which failed test was rerun after correction?
- What is backed up?
- Where does the backup live?
- How is restore tested without overwriting live records?
- What current-state summary would let someone resume the work?
- What decision history matters?
- What stop controls must be documented?
- Who could take over during an absence?
- What access arrangements are required?
- Which records should be portable to another AI tool?
- What portability gap remains?
- What does the workflow cost to operate, in minutes per month?
- What value does it return, in minutes and in the intake’s own terms?
- When are the 30-, 90-, 180-, and 365-day reviews scheduled? (Calendar all four now, each with a standing agenda: measures vs. baseline, failures and corrections, memory contradiction audit, artifact review dates, reinforcement runway status.)
- Where do the reinforcement runway touchpoints land? (Per the runway protocol in Part I: day 3, day 7, day 30, day 66, day 90, calendar each with its delayed-recall exercise. The day-30 runway checkpoint and the day-30 operating review are the same meeting.)
Build prompt
Using the interview answers, create an operating review and recovery plan.
Include: owner; review cadence; selected measures with intake baseline and
current numbers; observed improvements; failures and causes; corrections and
retests; backup scope and location; restore procedure (to a separate location,
never over live records); current-state resume instructions; decision history
location; stop controls; continuity handoff (who takes over, what access they
need); portability notes and the explicit remaining gap; operating cost and
value in minutes; review calendar (30/90/180/365 dates with standing agenda);
reinforcement runway touchpoints (day 3/7/30/66/90, each with its exercise and
a calendar date).
Do not claim portability, savings, recovery, or authority the evidence does not
establish. Separate approved facts / proposed rules / assumptions / unanswered
questions / evidence needed; mark assumptions unconfirmed. Numbers come from
the intake brief and the Evidence Log, or they are marked unconfirmed.Test
Concrete cases:
- Correction loop. Take one recorded failure from the Evidence Log, correct the underlying record or rule, rerun that exact case. Pass if it now passes and the correction is logged.
- Backup/restore. Restore the backup to a safe separate location (a test folder or second vault, never over live records). Pass if the restored set opens, essential records match the live set (spot-check workflow brief, decision record, action matrix, run log), and live records are verified untouched afterward.
- Fresh-session recovery. In a fresh session with only the maintained records, recover current state, key decision reasons, next action, and stop controls. Pass if all four are found without asking the owner.
- Portability run (mandatory at full depth). Load the canonical artifact set into a second suitable model or tool and run one representative task. Pass if it meets its quality checks there; if not, document the specific gap, which artifact failed and why, and add it to the review agenda.
Done
Definition of done
Essential records restore correctly; the fresh-session task meets its quality checks; a maintainer can find the next action and stop controls without the owner; the portability run is executed or its gap is explicit; the 30/90/180/365 reviews and runway touchpoints are all on a real calendar; the first “after” numbers are recorded against the intake baseline.
Minimum viable artifact
Review owner, measures with baseline, backup location, restore test record, continuity note, next review date, and the runway schedule.
Depth by stage
I10-lite (Expert): owner, cadence, backup scope and location, one restore test, continuity note, day-30/90 reviews calendared; portability documented as a gap, not yet run. I10-full (Frontier Operator): everything above plus the mandatory portability run, the full 30/90/180/365 calendar with standing agendas, the continuity handoff rehearsed with the named successor, and the runway through day 90. Graduation follows at full depth.
Going deeper.
Trust needs evidence
By this point the system runs: the Operator drafts, a human approves, the vault holds. The question that remains is the one most people never ask: how would we know this is working next month? Not “does it feel useful”: what would we actually look at?
Without measurement, two failure modes creep in quietly. Either you over-trust: the system drifts off-voice and off-doctrine and nobody notices until a tenant or a client does. Or you under-trust: the system is performing well but you keep re-doing its work, and the advantage you built goes unused. Measurement is how you earn calibrated trust: enough evidence to delegate more, early warning when something slips.
The test: how would we know this is working next month? If the answer is a shrug, this order is not installed.
The two files
Reading the signals
- On-voice without prompting. Count how often you rewrote a draft for tone this month. Falling is health; rising means Brand Voice needs sharper examples: fix the file, not the drafts.
- Cites authority correctly. Spot-check five material claims. Each should carry a level, and the level should be right. Wrong or missing levels point at Order 2.
- Escalations appropriate, not constant. Zero escalations means the system is guessing instead of asking. Constant escalation means scope or zones are drawn wrong. Both are findings about the limits, not Operator moods.
- Drafts trusted enough to use. The bluntest signal: what share of Draft & Wait output ships with light edits? If you quietly rewrite everything, the foundation is not working. Find which document is lying.
The correction loop
Measurement only matters if findings change the files. The loop is always the same: miss → trace → edit → retest. A tone miss traces to Brand Voice. An invented policy traces to a gap in doctrine. A wrong action traces to zones or scope. Three misses with the same root cause is a standing order to edit the document that caused them, and the promote/demote question keeps the vault honest: validated drafts move up, stale “doctrine” moves down or out.
Foundations decay by default
Everything you built under the first five orders is accurate today. It will not stay accurate on its own. Priorities shift, constraints change, people come and go, and a foundation nobody maintains quietly becomes fiction, while the Operator keeps citing it with full confidence. That is the failure mode this order exists to prevent: not a crash, but a slow divergence between the files and reality.
The fix is not effort. It is rhythm: every order has an owner, every review has a date, and every change has a protocol. Minutes per week, not hours.
The second test: who owns each order, and when is it reviewed? Two questions. If either draws a blank, the foundation is already decaying.
A cadence that fits a real calendar
- Weekly: about 15 minutes. Clear Session Notes: promote anything worth keeping, archive the rest. Skim the approval log for misses worth tracing.
- Monthly: about 30 minutes. Run the Measurement review: the four questions, one dated entry. Re-touch Strategic Priorities and Active Constraints if the cycle has turned. Edit any document with three same-cause misses.
- Quarterly: about an hour. Re-run the Scorecard and compare to last quarter. Reread the Identity Charter: still true? Audit zones and scope: has the Operator earned wider trust anywhere, or drifted anywhere? Verify succession access still works.
The update protocol
Changes to the foundation are Draft & Wait work, like everything else that matters:
- Proposed in writing. Anyone, including the Operator, can propose an edit. The Operator may draft it; it may never apply it. Foundation edits are always human-approved.
- Approved by the order's owner. The named human from the ownership table signs off before doctrine changes.
- Versioned simply. A dated changelog line at the top of the file, or the vault in a git repository if that is your world. What matters is that “what changed and when” has an answer.
Score yourself
The Readiness Scorecard is this order turned into a tool: a weighted self-assessment across all six orders that names your weakest order and what to fix first. Run it when you finish the build, then re-run it at quarterly reviews: the score trend is itself a measurement.
You are done with Section 4.2 when
- Measurement Notes exists and names the signals you actually check
- Operating Rhythm exists with a name on every ownership line
- Weekly, monthly, and quarterly reviews are on an actual calendar, and the first ones have happened
- At least one document has been edited because of a finding: the loop is real
- Foundation edits follow the protocol: proposed, approved, versioned
- You can answer “is this working?” with evidence instead of a feeling
The program, complete
That is Groundwork: six standing orders, fifteen files, one Operator, one loop. Know me, so it sounds like you. Know my information, so it trusts the right things. Know your limits, so it acts only where allowed. Help me operate: one scoped Operator doing real work. Remember, so it survives you. Improve, so it stays true. And from activation on, the loop underneath it all: capture, understand, decide, act, verify, remember, improve.
Two ways to close the loop: run the Readiness Scorecard to get your baseline score and your weakest order, or book the free Audit and walk through it with Adam. Questions: adam@adamabdalla.com.
The documents this section produces.
Document 12: Continuity & Succession
01_Identity/Continuity.md
How the foundation survives people leaving and systems failing. It lives in the Identity folder on purpose: what survives you is part of who you are.
# Continuity & Succession
## Key context that must survive
- Identity Charter
- Decision Principles
- Authority Classification
- Current Strategic Priorities
- Critical operating files
## Succession rules
- Who inherits decision rights:
- How access is transferred:
- What must be reviewed before transfer:
## Failure & override
- How the system is stopped:
- Who has final override authority:
- What gets preserved during failure:
How to write it well: use names, not roles-you-hope-to-hire. “Your brother gets read access and your attorney gets the override call” is a real succession rule. Then test the boring parts: does the successor actually have the passwords, and do they know the vault exists?
Document 14: Measurement Notes
06_Output_Standards/Measurement_Notes.md
How you will know the foundation is working, written down so the check actually happens.
# Measurement Notes
## Signals of a working foundation
- Output stays on-voice without prompting
- The system cites authority correctly
- Escalations are appropriate, not constant
- Humans trust the drafts enough to use them
## Review questions (monthly)
- What broke this month?
- What did the system get wrong?
- What context was missing?
- What should be promoted or demoted
in authority?
How to use it: the four signals are observations, not metrics dashboards. You can answer all four from a week of normal use plus your approval log. Write one dated entry per month under the questions. Six entries in, you have something rare: an evidence trail of whether your AI system is getting better or worse.
Document 15: Operating Rhythm
06_Output_Standards/Operating_Rhythm.md
Ownership, cadence, and how changes get made. The last document in the program, and the one that keeps the other fourteen alive.
# Operating Rhythm
## Ownership
- Know me (Identity):
- Know my information (Knowledge):
- Know your limits (Governance):
- Help me operate (the Operator):
- Remember (Continuity):
- Improve:
## Review cadence
- Weekly:
- Monthly:
- Quarterly:
## Update protocol
- How changes are proposed:
- Who approves:
- How versioning is handled:
How to write it well: in a small operation, every line may say the same name, and that is fine. The point is that it is written, so the day someone else joins, ownership transfers as an edit instead of an excavation.