Roadmap To Be A Data Engineer / Lesson 33
The Exercise Calendar
Why one big recovery drill a year proves almost nothing.
Thesis. A capability is not something a platform has. It is something it demonstrated recently, on a version of itself that still exists. Evidence expires at a rate set by how fast the thing underneath it changes; an exercise is the only instrument that produces such evidence; and the two dials — how often and how much — multiply, which is why the answer is small and frequent rather than large and annual.
1The scene
Last September the company ran a full-scale exercise (Lesson 30): six objectives, one performed without challenges, an eleven-item improvement plan, and the line every such report carries — repeat this exercise annually. Somebody put 18 September 2026 in the calendar.
Twelve months later the audit committee asks the question the programme was never designed to answer: take a Tuesday at random from the last year — could you have restored the dropped table, rebuilt the revenue report, reproduced the figure approved in July? Not do you think so: could you have shown it?
On 357 of 365 days, not one of the eight things the platform is supposed to be able to do was in a state anyone could demonstrate. Not because the platform got worse — because it changed 320 times and the evidence was collected once.
2Currency
Aviation solved this with a word. 14 CFR 61.57: “at least three takeoffs and three landings within the preceding 90 days”; for instrument flight, six approaches, holding procedures and tasks, and intercepting and tracking courses “within the 6 calendar months preceding the month of the flight”, with only a full proficiency check to re-establish it after a further six. Three things worth stealing: the evidence expires, it is scoped (per aircraft category and class), and there is a grace period after which re-establishing costs more than maintaining would have.
What cannot be stolen is the number. Ninety days works because the aeroplane in April is the same aeroplane it was in January. Over the observed year this platform absorbed 224 model merges, 62 config changes, 11 new copies of the data, 9 rota changes and 14 vendor-side changes. A calendar-based currency rule is measuring a clock that has nothing to do with the thing it certifies.
| Capability | From | Invalidations/yr | Mean days between | |
|---|---|---|---|---|
| C1 | Restore a table somebody dropped | L30 obj.1 | 21 | 17.38 |
| C2 | Rebuild every served object from source | L30 obj.3 | 249 | 1.47 |
| C3 | Reproduce a figure published last quarter | L32 | 170 | 2.15 |
| C4 | Notice a materially wrong number unprompted | L25 / L31 | 50 | 7.30 |
| C5 | Replay a source from the landing zone | L22 / L30 obj.4 | 19 | 19.21 |
| C6 | Get a named defect to the person who can fix it | L28 | 16 | 22.81 |
| C7 | Erase a data subject from every copy | L14 / L24 | 29 | 12.59 |
| C8 | Say who could have seen a customer’s data | L23 | 51 | 7.16 |
A 15.6× spread in the same company in the same year — rebuilding everything survives 1.47 days, knowing who to call survives 22.8 — and no exercise programme in general use distinguishes between them.
3The report was stale before it was circulated
The exercise of 18 Sep 2025 demonstrated four capabilities. Here is when each stopped being evidence:
| Lapsed on | Days stale when the AAR went out | Because | |
|---|---|---|---|
| C2 rebuild-all | 2025-09-20 | 19 | a merge touched a model on the rebuild path |
| C5 replay-source | 2025-09-22 | 17 | a retention setting changed on a source |
| C3 reproduce | 2025-09-25 | 14 | a merge touched the tier-1 lineage |
| C1 restore-table | 2025-09-26 | 13 | a partitioning setting changed |
The after-action report was circulated on 2025-10-09, 21 days after the exercise — which is fast. By then all four of the capabilities it certified had lapsed, the rebuild by 19 days. The document describing what the company could do was, on the day it was published, a description of a platform that no longer existed.
Across the whole year the annual programme holds 21 capability-days out of 2,920. Four capabilities were never exercised at all.
The zero is arithmetic, not luck. With 21 capability-days, 120 incidents over 365 days and 47 of them needing one of these capabilities, the expected number of incidents the annual exercise could have helped with is 0.48. The realised number is 0.
4The coverage law
coverage(λ, T) = (1 − e^(−λT/365)) / (λT/365)
The expected value of min(time to the next invalidating change, T) divided by T. Coverage depends on λ and T only through their product, so halving the scope of an exercise is worth exactly as much as doubling its frequency. Validated against the generated change stream at five cadences × eight capabilities, each averaged over every start offset: max deviation 7.1 pp, and the simulation is always slightly above the formula because real merges arrive more regularly than a Poisson process.
To be current half the time:
| Capability | λ/yr | Exercise every (days) | Runs a year | |
|---|---|---|---|---|
| C1 | restore-table | 21 | 27.7 | 13 |
| C2 | rebuild-all | 249 | 2.3 | 156 |
| C3 | reproduce | 170 | 3.4 | 107 |
| C4 | detect-wrong | 50 | 11.6 | 31 |
| C5 | replay-source | 19 | 30.6 | 12 |
| C6 | route-repair | 16 | 36.4 | 10 |
| C7 | erase-subject | 29 | 20.1 | 18 |
| C8 | prove-access | 51 | 11.4 | 32 |
The full-scale exercise costs €6,048 a run. Running it every 2.34 days — what being half-current on rebuild everything costs — is 156 exercises a year and €944,986. An exercise performed by people in a room cannot be run often enough to keep evidence fresh, at any budget a company would recognise.
5The calendar invite deletes the incident
Lesson 28 measured where a real incident’s 189 h elapsed time goes: 52 h finding the owning team, 97 h waiting to be picked up, 14.7 h of actual work, 26 h waiting for a release. A scheduled exercise deletes routing (the invitation names the scenario), waiting (the people are in the room) and the release train (the fix goes in a sandbox).
| Rehearsal | Measures | Share of the clock | Understates by |
|---|---|---|---|
| announced | 14.7 h | 7.8% | 12.82× |
| unannounced | 162.9 h | 86.2% | 1.16× |
| real | 189.0 h | 100.0% | 1.00× |
An exercise in the calendar measures 7.8% of the clock and reports success. The fix is not to make every drill unannounced — two unannounced injections a year cost €864, of which 7 of 12 hours is real work other people dropped. Run the cheap announced drills constantly and two unannounced ones a year purely to measure what the announced ones under-report by, then quote every announced drill with its factor attached.
6Six programmes, costed
Damage is measured in object-hours wrong — one served answer wrong for one hour — over the same 120 incidents.
| Programme | €/yr | Person-hours | Capabilities current (mean of 8) | Wrong-answer hours removed | € per 1,000 h removed | |
|---|---|---|---|---|---|---|
| P0 | no exercise of any kind | €0 | 0 | 0.00 | 0.00% | — |
| P1 | the full-scale exercise, once a year | €6,048 | 84 | 0.06 | 0.00% | — |
| P3 | read-through ×3, workshop, quarterly tabletop, annual full-scale | €12,096 | 168 | 0.06 | 0.00% | — |
| P2 | the full-scale exercise, quarterly | €30,240 | 420 | 0.37 | 0.02% | 1,374,545 |
| P5 | seven automated drills, no meetings | €5,851 | 70 | 3.27 | 35.92% | 152 |
| P4 | P5 + two unannounced injections + the annual full-scale | €13,195 | 172 | 3.39 | 39.25% | 314 |
- The obvious fix loses. Quarterly is five times the money for 0.02% — twenty-two object-hours — at €1,374,545 per thousand removed. 91 days and 2.3 days are not close.
- The framework programme removes nothing, by definition. HSEEP’s discussion-based types (seminar, workshop, tabletop, game) verify the plan; a tabletop is “a discussion-based exercise … intended to generate a dialogue”. Nothing is run, so no capability becomes current. Three of the twelve catalogued exercises are discussion-based and all three demonstrate zero capabilities — which is what they are for; the criticism is of a programme made only of them, at €12,096.
- The machines are cheaper than the meeting. P5 costs €196.69 less than running the full-scale exercise once and removes 35.92% instead of 0.00%. Compute for all seven automated drills is €787.31/yr, of which €609 is one licence.
Where the reduction comes from — the three levers, turned on one at a time:
| Lever | Of the 39.25 pp |
|---|---|
| Detection — the canary and the injection | 35.43 pp |
| Routing and waiting — a rehearsed page path, a named owner | 3.33 pp |
| The repair itself — restore, rebuild, replay, re-derive | 0.48 pp |
The thing everyone means by “DR testing” is worth 0.48 percentage points. The thing nobody calls an exercise — injecting a known defect and confirming something notices — is worth 35.43. Lesson 25 restated: almost all the damage accrues before anyone knows.
Robustness. Six repair durations are inputs, so the table was re-run with the cost of being unrehearsed scaled down:
| Penalty kept | Baseline (object-hours) | P4 removes | P5 removes |
|---|---|---|---|
| 100% | 107,092 | 39.25% | 35.92% |
| 50% | 105,874 | 39.46% | 36.08% |
| 25% | 105,266 | 39.56% | 36.17% |
| 0% | 104,657 | 39.67% | 36.26% |
At zero — where rehearsing a repair is worth literally nothing — the programme still removes 39.67%. The recommendation does not rest on any of the invented numbers.
7Allocation, and discovery
One served object a night, eighteen of them. The metric you would naturally report ranks the options backwards:
| Allocation | Mean coverage, all 18 | Coverage of the 8 tier-1 | Object-hours wrong |
|---|---|---|---|
| one slot each (18-night cycle) | 41.83% | 34.62% | 1,299 |
| three slots for tier-1 (34-night cycle) | 37.02% | 49.80% | 1,100 |
| tier-1 only (8-night cycle) | 27.79% | 62.53% | 850 |
Covering everything equally gives the best headline (41.83%) and the worst outcome; tier-1 only gives the worst headline (27.79%) and 34.6% less damage, because those eight carry 50.6% of it.
Currency and discovery are two products billed under one name. Currency is the freshness of the evidence; discovery is finding a gap that was already there. Currency wants the same scope often; discovery wants new scope rarely.
Assume a change to a recovery path leaves it broken 4% of the time. Across the eighteen rebuild paths the year opened 32 gaps. The annual exercise closed 0 of them inside the year (6,182 gap-days open); the rotating nightly drill found all 32 at a mean age of 8.7 days (276 gap-days) — 22.4× less exposure.
What a rotation cannot cover is the scope that only exists when everything runs at once: 2 of 6 of last year’s objectives (rebuild timing under contention, Lesson 27; six-team coordination, Lesson 28). That is the case for keeping the full-scale exercise — once a year, for two of its six objectives. The marginal €7,344 P4 spends over P5 buys 3,571 object-hours at €2.06 each against €0.15 for the automated part — 13.5× worse per euro, and still worth buying.
8The Tuesday sandbox: two silent defaults
BigQuery documentation: “A table clone is a lightweight, writable copy of another table” … “You are only charged for storage of data in the table clone that differs from the base table, so initially there is no storage cost for a table clone.” But the limitations matter: you cannot clone further back than the time travel window, which Lesson 30 found is seven days by default in four independent places here.
dbt clone‘s behaviour comes from three macros in dbt-adapters:
{% macro default__can_clone_table() %}
{{ return(False) }}
{% endmacro %}
{# clone.sql: #}
-- If this is a database that can do zero-copy cloning of tables, and the other
-- relation is a table, then this will be a table
-- Otherwise, this will be a view
dbt-bigquery overrides it to True; dbt-duckdb, like any adapter that has not implemented it, takes the default. Run against a three-model project (dbt-core 1.12.5, dbt-duckdb 1.11.0):
$ dbt clone --state state --target drill
Completed successfully
Done. PASS=1 WARN=0 ERROR=0 SKIP=0 NO-OP=0 REUSED=0 TOTAL=1
>>> select table_schema, table_name, table_type from information_schema.tables
[('drill', 'fct_orders', 'VIEW'), ('main', 'fct_orders', 'BASE TABLE')]
>>> delete from drill.fct_orders where order_id = 1
BinderException: Can only delete from base table
$ dbt clone --state state --target drill # after production changed
Relation "prod"."drill"."fct_orders" already exists
Done. PASS=1 WARN=0 ERROR=0 SKIP=0 NO-OP=0 REUSED=0 TOTAL=1
The sandbox is a view onto production, the drill you most want to run cannot be run in it, and on an adapter that can clone the opposite bites: the clone is a frozen copy and a re-run without --full-refresh is a no-op, so the Tuesday drill runs against whatever production looked like when the sandbox was first created. Same command, two adapters, two opposite silent defaults, PASS=1 in both cases. A drill environment needs its own assertion before the drill: that the sandbox is a table and not a view, and that its newest row is from today.
9The instrument: capabilities current today
An integer between 0 and 8. No new data — the last successful run of each drill is in the CI artifacts, the last change to each path is in git and the config history. A config read, not a data read; €0.00 a year.
with last_run as (
select d.day_ix as t, r.cap_id, max(r.day_ix) as demonstrated
from days d join exercise_runs r on r.day_ix <= d.day_ix
group by 1, 2),
struck as (
select l.t, l.cap_id, l.demonstrated,
count(i.day_ix) filter (
where i.day_ix > l.demonstrated and i.day_ix <= l.t) as invalidated_since
from last_run l
left join cap_invalidations i on i.cap_id = l.cap_id
group by 1, 2, 3)
select t, count(*) filter (where invalidated_since = 0) as capabilities_current
from struck group by t order by t
| Programme | Mean | Days reading 0 | Days ≥ 4 | Best day | |
|---|---|---|---|---|---|
| P0 | no exercise of any kind | 0.00 | 365 | 0 | 0 |
| P1 | the full-scale exercise, once a year | 0.06 | 357 | 2 | 4 |
| P3 | read-through ×3, workshop, quarterly tabletop, annual full-scale | 0.06 | 357 | 2 | 4 |
| P2 | the full-scale exercise, quarterly | 0.37 | 292 | 8 | 4 |
| P5 | seven automated drills, no meetings | 3.27 | 0 | 156 | 6 |
| P4 | P5 + two unannounced injections + the annual full-scale | 3.39 | 0 | 168 | 8 |
A canyon, not a threshold. It reports that the evidence is fresh, never that the capability is adequate — a drill scoped to the wrong object, or one that passes because it was announced, lights the same cell as a good one.
10What no exercise reaches
Scoring the twenty-one failure shapes from Lessons 06–32: a recurring exercise reaches 10 outright and 3 in part; a test suite (Lesson 31) reaches 1 outright and 3 in part; 2 shapes are reached by both and 6 by neither.
| Shape | Exercise | By | Test suite | Why |
|---|---|---|---|---|
| 06-10 event | part | E9 | yes | a monitor catches it; the drill only proves the monitor fires |
| 11 constant | no | — | part | consistent across every rebuild, so a rebuild-and-diff sees nothing |
| 12 drift | no | — | part | each increment is inside the band; only the sum is not |
| 13 unwatched | yes | E11 | no | a copy-map walk crosses the boundary nothing instruments |
| 14 reversible | yes | E11 | no | erase a seeded subject, let the nightly refresh run, look again |
| 15 unreproduced | yes | E5 | no | the drill is the production-shaped environment the tests lacked |
| 16 bundled | no | — | no | a decision with one name on it; no exercise has an opinion |
| 17 transient | no | — | no | 9.67 s a night: no sampler reaches it, exercises least of all |
| 18 extremal | part | E8 | no | only a rebuild at production scale meets the hot key |
| 19 referential | no | — | no | the error is in a correspondence no table stores |
| 20 plural | no | — | no | several right answers; an exercise has nothing to assert |
| 21 remote | no | — | no | knowable statically before anything ran; a build step, not a drill |
| 22 absent | yes | E6 | no | replay the source and reconcile: the second count from outside |
| 23 permitted | yes | E12 | no | re-run the access ladder against the IAM state of the day |
| 24 detached | yes | E11 | no | the copy map is the drill’s inventory |
| 26 unsettled | no | — | no | nothing is wrong; the quantity has not finished happening |
| 27 contended | part | E8 | no | the interleaving only exists when everything runs at once |
| 28 unclaimed | yes | E9 | no | the unannounced injection measures the handoff queue directly |
| 29 anachronistic | yes | E7 | part | a replay that withholds the future is exactly this exercise |
| 30 irrecoverable | yes | E5 | no | the reproducible horizon is only ever measured by rebuilding |
| 32 unprovable | yes | E7 | no | re-derive a published figure and diff it against what was said |
The two instruments are nearly disjoint — they are not competing budgets and not substitutes. The residue (plural, referential, unsettled, bundled, remote, drift) has never had a better instrument as its fix; it needs a representational change. So an honest assurance argument has three parts: tests for the predicates that can be written, exercises for the capabilities that can only be demonstrated, and architecture for the rest — with an explicit written list of what is in the third bucket. The first two are budget lines; the third is a roadmap.
11What to ask the team
- “When did we last demonstrate this, and what has changed since?” A date on its own is a calendar answer to a condition question.
- “Was the drill announced?” If yes, multiply the reported time by the factor and say so. If nobody knows the factor, the next thing to schedule is the unannounced drill that measures it.
- “What is the drill environment, exactly — a clone or a view?” And when was it last refreshed.
- “Which of the eight could we not do today, and is that the plan?” Not being current is a legitimate choice; not knowing is not.
- “Which objectives genuinely need everything running at once?” That list, and only that list, is what the annual exercise is for.
12Hands-on, twenty minutes
- Write down eight capabilities as “we can verb a thing“, each one something a real incident has needed.
- Put a date on each: when did anyone last do it end to end on production data?
- Count λ:
git log --since=1.year --oneline -- path/to/its/models | wc -l, plus the config changes. 365 ÷ that is how long your evidence lasts. - Compute coverage with the § 33.04 formula — one line of arithmetic per capability.
- Automate the cheapest one (almost always “rebuild one thing into a scratch schema and diff it”) on a nightly rotation over your served objects, tier-1 first. Assert the sandbox is a table and not a view before it runs.
- Publish the integer. Let it be embarrassing for a quarter.
13Takeaway and vocabulary
- An annual exercise buys a document, not a capability. 21 capability-days out of 2,920, and the report was stale before it was circulated.
- A drill in the calendar has already done the slow part of the incident. It measures 7.8% of the clock. Run two unannounced ones a year purely to learn the factor.
- Rehearsing the repair is worth half a point; rehearsing the detection is worth thirty-five. Spend accordingly.
| Term | Meaning |
|---|---|
| Currency | Whether a capability has been demonstrated recently enough for the demonstration still to be evidence. |
| Verified-day coverage | The share of the year in which a capability is current — what an exercise programme actually buys. |
| Invalidating change | A change to anything a capability depends on. Its rate λ is countable from git and the config history. |
| Discussion-based vs operations-based | HSEEP’s own split. Only the second kind can make a capability current. |
| Announced-drill factor | How much a scheduled rehearsal understates the real event. Measured, not assumed. |
| Gap | A recovery path that is broken and untried. Gap-days, not gap count, is the exposure. |
| Discovery vs currency | The two products an exercise sells; they want opposite cadences. |
| Steady-state hypothesis | From the chaos principles: “focus on the measurable output of a system, rather than internal attributes.” For a data platform the measurable output is an answer. |
14How these numbers were made
One deterministic generator (no RNG, no seed, no wall-clock input) over a synthetic year on a platform shaped like Lesson 30’s: 234 objects, 452 edges, 18 served answers of which 8 are tier-1, 11 external systems, 120 incidents.
Counted: every λ, every currency series, every coverage figure, the gap-days, the allocation table, the instrument’s distribution.
Carried in from earlier lessons, where they were measured: the four phase shares and the 189 h mean (L28); the 1.33 h alert and 16.5-day human detection lags and the 120-incident population (L25/L28); the 4 min 12 s restore and the 20.06 h / 69.95 h rebuild pair (L30); the €609 mutation-canary licence (L31).
Inputs: the change rates; €72/h loaded; build and upkeep hours for the automated drills; six unrehearsed repair durations; the 4% gap probability; the 21-day AAR lag. The repair durations are load-bearing, which is why the robustness table sets them to zero.
Checks. The closed form was validated against the simulation with phase-averaging (an earlier version compared a single annual draw against an expectation and reported up to 4 pp of pure sampling noise). The 0.00% was attacked as a round number before publication and survived with its expectation published beside it. The published SQL was re-run in DuckDB: 125 checks, 125 matched, 0 failed — one disagreement found on the way, Python’s sorted(x)[len(x)//2] against SQL’s quantile_cont(0.5) on an even-length list (the Lesson 18 median trap). dbt clone was actually run rather than described. Sources read at source: 14 CFR 61.57 (eCFR), the HSEEP doctrine 2020 revision, the chaos-engineering principles, the BigQuery table-clones documentation, and the clone macros from the shipped dbt-adapters 1.24.5 and dbt-bigquery 1.12.1 wheels. Round trip: both source files extracted from the rendered page, run in a clean directory, 95 top-level keys / 755 scalar values, 0 differences, figures.json md5-identical across runs and directories.
Two planned claims lost. The lesson was drafted to argue that unannounced drills should replace announced ones — costing the interruption showed they are a tax on the attention being measured, and the surviving recommendation is two a year used only to calibrate the rest. It was also drafted to argue for weighting the nightly rotation towards tier-1; that is right, but for the opposite reason to the one planned — the allocation with the worst mean coverage is the best one, and the metric ranks the three options exactly backwards.
No new failure shape. Like Lessons 25 and 31 this is a measurement of an instrument, not of the data. The scoring in § 33.10 is a judgement, laid out row by row so it can be argued with.
Artifact (visual version with six figures): published 2026-09-21.