Roadmap To Be A Data Engineer / Lesson 33

Lesson 33 Game days Fundamentals §6 Data Quality / §5 Orchestration (+§7) About 20 min read

The Exercise Calendar

Why one big recovery drill a year proves almost nothing.

Thesis. A capability is not something a platform has. It is something it demonstrated recently, on a version of itself that still exists. Evidence expires at a rate set by how fast the thing underneath it changes; an exercise is the only instrument that produces such evidence; and the two dials — how often and how much — multiply, which is why the answer is small and frequent rather than large and annual.

1The scene

Last September the company ran a full-scale exercise (Lesson 30): six objectives, one performed without challenges, an eleven-item improvement plan, and the line every such report carries — repeat this exercise annually. Somebody put 18 September 2026 in the calendar.

Twelve months later the audit committee asks the question the programme was never designed to answer: take a Tuesday at random from the last year — could you have restored the dropped table, rebuilt the revenue report, reproduced the figure approved in July? Not do you think so: could you have shown it?

On 357 of 365 days, not one of the eight things the platform is supposed to be able to do was in a state anyone could demonstrate. Not because the platform got worse — because it changed 320 times and the evidence was collected once.

2Currency

Aviation solved this with a word. 14 CFR 61.57: “at least three takeoffs and three landings within the preceding 90 days”; for instrument flight, six approaches, holding procedures and tasks, and intercepting and tracking courses “within the 6 calendar months preceding the month of the flight”, with only a full proficiency check to re-establish it after a further six. Three things worth stealing: the evidence expires, it is scoped (per aircraft category and class), and there is a grace period after which re-establishing costs more than maintaining would have.

What cannot be stolen is the number. Ninety days works because the aeroplane in April is the same aeroplane it was in January. Over the observed year this platform absorbed 224 model merges, 62 config changes, 11 new copies of the data, 9 rota changes and 14 vendor-side changes. A calendar-based currency rule is measuring a clock that has nothing to do with the thing it certifies.

Capability From Invalidations/yr Mean days between
C1 Restore a table somebody dropped L30 obj.1 21 17.38
C2 Rebuild every served object from source L30 obj.3 249 1.47
C3 Reproduce a figure published last quarter L32 170 2.15
C4 Notice a materially wrong number unprompted L25 / L31 50 7.30
C5 Replay a source from the landing zone L22 / L30 obj.4 19 19.21
C6 Get a named defect to the person who can fix it L28 16 22.81
C7 Erase a data subject from every copy L14 / L24 29 12.59
C8 Say who could have seen a customer’s data L23 51 7.16

A 15.6× spread in the same company in the same year — rebuilding everything survives 1.47 days, knowing who to call survives 22.8 — and no exercise programme in general use distinguishes between them.

3The report was stale before it was circulated

The exercise of 18 Sep 2025 demonstrated four capabilities. Here is when each stopped being evidence:

Lapsed on Days stale when the AAR went out Because
C2 rebuild-all 2025-09-20 19 a merge touched a model on the rebuild path
C5 replay-source 2025-09-22 17 a retention setting changed on a source
C3 reproduce 2025-09-25 14 a merge touched the tier-1 lineage
C1 restore-table 2025-09-26 13 a partitioning setting changed

The after-action report was circulated on 2025-10-09, 21 days after the exercise — which is fast. By then all four of the capabilities it certified had lapsed, the rebuild by 19 days. The document describing what the company could do was, on the day it was published, a description of a platform that no longer existed.

Across the whole year the annual programme holds 21 capability-days out of 2,920. Four capabilities were never exercised at all.

The zero is arithmetic, not luck. With 21 capability-days, 120 incidents over 365 days and 47 of them needing one of these capabilities, the expected number of incidents the annual exercise could have helped with is 0.48. The realised number is 0.

4The coverage law

coverage(λ, T)  =  (1 − e^(−λT/365)) / (λT/365)

The expected value of min(time to the next invalidating change, T) divided by T. Coverage depends on λ and T only through their product, so halving the scope of an exercise is worth exactly as much as doubling its frequency. Validated against the generated change stream at five cadences × eight capabilities, each averaged over every start offset: max deviation 7.1 pp, and the simulation is always slightly above the formula because real merges arrive more regularly than a Poisson process.

To be current half the time:

Capability λ/yr Exercise every (days) Runs a year
C1 restore-table 21 27.7 13
C2 rebuild-all 249 2.3 156
C3 reproduce 170 3.4 107
C4 detect-wrong 50 11.6 31
C5 replay-source 19 30.6 12
C6 route-repair 16 36.4 10
C7 erase-subject 29 20.1 18
C8 prove-access 51 11.4 32

The full-scale exercise costs €6,048 a run. Running it every 2.34 days — what being half-current on rebuild everything costs — is 156 exercises a year and €944,986. An exercise performed by people in a room cannot be run often enough to keep evidence fresh, at any budget a company would recognise.

5The calendar invite deletes the incident

Lesson 28 measured where a real incident’s 189 h elapsed time goes: 52 h finding the owning team, 97 h waiting to be picked up, 14.7 h of actual work, 26 h waiting for a release. A scheduled exercise deletes routing (the invitation names the scenario), waiting (the people are in the room) and the release train (the fix goes in a sandbox).

Rehearsal Measures Share of the clock Understates by
announced 14.7 h 7.8% 12.82×
unannounced 162.9 h 86.2% 1.16×
real 189.0 h 100.0% 1.00×

An exercise in the calendar measures 7.8% of the clock and reports success. The fix is not to make every drill unannounced — two unannounced injections a year cost €864, of which 7 of 12 hours is real work other people dropped. Run the cheap announced drills constantly and two unannounced ones a year purely to measure what the announced ones under-report by, then quote every announced drill with its factor attached.

6Six programmes, costed

Damage is measured in object-hours wrong — one served answer wrong for one hour — over the same 120 incidents.

Programme €/yr Person-hours Capabilities current (mean of 8) Wrong-answer hours removed € per 1,000 h removed
P0 no exercise of any kind €0 0 0.00 0.00% —
P1 the full-scale exercise, once a year €6,048 84 0.06 0.00% —
P3 read-through ×3, workshop, quarterly tabletop, annual full-scale €12,096 168 0.06 0.00% —
P2 the full-scale exercise, quarterly €30,240 420 0.37 0.02% 1,374,545
P5 seven automated drills, no meetings €5,851 70 3.27 35.92% 152
P4 P5 + two unannounced injections + the annual full-scale €13,195 172 3.39 39.25% 314
  • The obvious fix loses. Quarterly is five times the money for 0.02% — twenty-two object-hours — at €1,374,545 per thousand removed. 91 days and 2.3 days are not close.
  • The framework programme removes nothing, by definition. HSEEP’s discussion-based types (seminar, workshop, tabletop, game) verify the plan; a tabletop is “a discussion-based exercise … intended to generate a dialogue”. Nothing is run, so no capability becomes current. Three of the twelve catalogued exercises are discussion-based and all three demonstrate zero capabilities — which is what they are for; the criticism is of a programme made only of them, at €12,096.
  • The machines are cheaper than the meeting. P5 costs €196.69 less than running the full-scale exercise once and removes 35.92% instead of 0.00%. Compute for all seven automated drills is €787.31/yr, of which €609 is one licence.

Where the reduction comes from — the three levers, turned on one at a time:

Lever Of the 39.25 pp
Detection — the canary and the injection 35.43 pp
Routing and waiting — a rehearsed page path, a named owner 3.33 pp
The repair itself — restore, rebuild, replay, re-derive 0.48 pp

The thing everyone means by “DR testing” is worth 0.48 percentage points. The thing nobody calls an exercise — injecting a known defect and confirming something notices — is worth 35.43. Lesson 25 restated: almost all the damage accrues before anyone knows.

Robustness. Six repair durations are inputs, so the table was re-run with the cost of being unrehearsed scaled down:

Penalty kept Baseline (object-hours) P4 removes P5 removes
100% 107,092 39.25% 35.92%
50% 105,874 39.46% 36.08%
25% 105,266 39.56% 36.17%
0% 104,657 39.67% 36.26%

At zero — where rehearsing a repair is worth literally nothing — the programme still removes 39.67%. The recommendation does not rest on any of the invented numbers.

7Allocation, and discovery

One served object a night, eighteen of them. The metric you would naturally report ranks the options backwards:

Allocation Mean coverage, all 18 Coverage of the 8 tier-1 Object-hours wrong
one slot each (18-night cycle) 41.83% 34.62% 1,299
three slots for tier-1 (34-night cycle) 37.02% 49.80% 1,100
tier-1 only (8-night cycle) 27.79% 62.53% 850

Covering everything equally gives the best headline (41.83%) and the worst outcome; tier-1 only gives the worst headline (27.79%) and 34.6% less damage, because those eight carry 50.6% of it.

Currency and discovery are two products billed under one name. Currency is the freshness of the evidence; discovery is finding a gap that was already there. Currency wants the same scope often; discovery wants new scope rarely.

Assume a change to a recovery path leaves it broken 4% of the time. Across the eighteen rebuild paths the year opened 32 gaps. The annual exercise closed 0 of them inside the year (6,182 gap-days open); the rotating nightly drill found all 32 at a mean age of 8.7 days (276 gap-days) — 22.4× less exposure.

What a rotation cannot cover is the scope that only exists when everything runs at once: 2 of 6 of last year’s objectives (rebuild timing under contention, Lesson 27; six-team coordination, Lesson 28). That is the case for keeping the full-scale exercise — once a year, for two of its six objectives. The marginal €7,344 P4 spends over P5 buys 3,571 object-hours at €2.06 each against €0.15 for the automated part — 13.5× worse per euro, and still worth buying.

8The Tuesday sandbox: two silent defaults

BigQuery documentation: “A table clone is a lightweight, writable copy of another table” … “You are only charged for storage of data in the table clone that differs from the base table, so initially there is no storage cost for a table clone.” But the limitations matter: you cannot clone further back than the time travel window, which Lesson 30 found is seven days by default in four independent places here.

dbt clone‘s behaviour comes from three macros in dbt-adapters:

{% macro default__can_clone_table() %}
    {{ return(False) }}
{% endmacro %}

{# clone.sql: #}
-- If this is a database that can do zero-copy cloning of tables, and the other
-- relation is a table, then this will be a table
-- Otherwise, this will be a view

dbt-bigquery overrides it to True; dbt-duckdb, like any adapter that has not implemented it, takes the default. Run against a three-model project (dbt-core 1.12.5, dbt-duckdb 1.11.0):

$ dbt clone --state state --target drill
Completed successfully
Done. PASS=1 WARN=0 ERROR=0 SKIP=0 NO-OP=0 REUSED=0 TOTAL=1

>>> select table_schema, table_name, table_type from information_schema.tables
[('drill', 'fct_orders', 'VIEW'), ('main', 'fct_orders', 'BASE TABLE')]

>>> delete from drill.fct_orders where order_id = 1
BinderException: Can only delete from base table

$ dbt clone --state state --target drill      # after production changed
Relation "prod"."drill"."fct_orders" already exists
Done. PASS=1 WARN=0 ERROR=0 SKIP=0 NO-OP=0 REUSED=0 TOTAL=1

The sandbox is a view onto production, the drill you most want to run cannot be run in it, and on an adapter that can clone the opposite bites: the clone is a frozen copy and a re-run without --full-refresh is a no-op, so the Tuesday drill runs against whatever production looked like when the sandbox was first created. Same command, two adapters, two opposite silent defaults, PASS=1 in both cases. A drill environment needs its own assertion before the drill: that the sandbox is a table and not a view, and that its newest row is from today.

9The instrument: capabilities current today

An integer between 0 and 8. No new data — the last successful run of each drill is in the CI artifacts, the last change to each path is in git and the config history. A config read, not a data read; €0.00 a year.

with last_run as (
  select d.day_ix as t, r.cap_id, max(r.day_ix) as demonstrated
  from days d join exercise_runs r on r.day_ix <= d.day_ix
  group by 1, 2),
struck as (
  select l.t, l.cap_id, l.demonstrated,
         count(i.day_ix) filter (
           where i.day_ix > l.demonstrated and i.day_ix <= l.t) as invalidated_since
  from last_run l
  left join cap_invalidations i on i.cap_id = l.cap_id
  group by 1, 2, 3)
select t, count(*) filter (where invalidated_since = 0) as capabilities_current
from struck group by t order by t
Programme Mean Days reading 0 Days ≥ 4 Best day
P0 no exercise of any kind 0.00 365 0 0
P1 the full-scale exercise, once a year 0.06 357 2 4
P3 read-through ×3, workshop, quarterly tabletop, annual full-scale 0.06 357 2 4
P2 the full-scale exercise, quarterly 0.37 292 8 4
P5 seven automated drills, no meetings 3.27 0 156 6
P4 P5 + two unannounced injections + the annual full-scale 3.39 0 168 8

A canyon, not a threshold. It reports that the evidence is fresh, never that the capability is adequate — a drill scoped to the wrong object, or one that passes because it was announced, lights the same cell as a good one.

10What no exercise reaches

Scoring the twenty-one failure shapes from Lessons 06–32: a recurring exercise reaches 10 outright and 3 in part; a test suite (Lesson 31) reaches 1 outright and 3 in part; 2 shapes are reached by both and 6 by neither.

Shape Exercise By Test suite Why
06-10 event part E9 yes a monitor catches it; the drill only proves the monitor fires
11 constant no — part consistent across every rebuild, so a rebuild-and-diff sees nothing
12 drift no — part each increment is inside the band; only the sum is not
13 unwatched yes E11 no a copy-map walk crosses the boundary nothing instruments
14 reversible yes E11 no erase a seeded subject, let the nightly refresh run, look again
15 unreproduced yes E5 no the drill is the production-shaped environment the tests lacked
16 bundled no — no a decision with one name on it; no exercise has an opinion
17 transient no — no 9.67 s a night: no sampler reaches it, exercises least of all
18 extremal part E8 no only a rebuild at production scale meets the hot key
19 referential no — no the error is in a correspondence no table stores
20 plural no — no several right answers; an exercise has nothing to assert
21 remote no — no knowable statically before anything ran; a build step, not a drill
22 absent yes E6 no replay the source and reconcile: the second count from outside
23 permitted yes E12 no re-run the access ladder against the IAM state of the day
24 detached yes E11 no the copy map is the drill’s inventory
26 unsettled no — no nothing is wrong; the quantity has not finished happening
27 contended part E8 no the interleaving only exists when everything runs at once
28 unclaimed yes E9 no the unannounced injection measures the handoff queue directly
29 anachronistic yes E7 part a replay that withholds the future is exactly this exercise
30 irrecoverable yes E5 no the reproducible horizon is only ever measured by rebuilding
32 unprovable yes E7 no re-derive a published figure and diff it against what was said

The two instruments are nearly disjoint — they are not competing budgets and not substitutes. The residue (plural, referential, unsettled, bundled, remote, drift) has never had a better instrument as its fix; it needs a representational change. So an honest assurance argument has three parts: tests for the predicates that can be written, exercises for the capabilities that can only be demonstrated, and architecture for the rest — with an explicit written list of what is in the third bucket. The first two are budget lines; the third is a roadmap.

11What to ask the team

  1. “When did we last demonstrate this, and what has changed since?” A date on its own is a calendar answer to a condition question.
  2. “Was the drill announced?” If yes, multiply the reported time by the factor and say so. If nobody knows the factor, the next thing to schedule is the unannounced drill that measures it.
  3. “What is the drill environment, exactly — a clone or a view?” And when was it last refreshed.
  4. “Which of the eight could we not do today, and is that the plan?” Not being current is a legitimate choice; not knowing is not.
  5. “Which objectives genuinely need everything running at once?” That list, and only that list, is what the annual exercise is for.

12Hands-on, twenty minutes

  1. Write down eight capabilities as “we can verb a thing“, each one something a real incident has needed.
  2. Put a date on each: when did anyone last do it end to end on production data?
  3. Count λ: git log --since=1.year --oneline -- path/to/its/models | wc -l, plus the config changes. 365 ÷ that is how long your evidence lasts.
  4. Compute coverage with the § 33.04 formula — one line of arithmetic per capability.
  5. Automate the cheapest one (almost always “rebuild one thing into a scratch schema and diff it”) on a nightly rotation over your served objects, tier-1 first. Assert the sandbox is a table and not a view before it runs.
  6. Publish the integer. Let it be embarrassing for a quarter.

13Takeaway and vocabulary

  • An annual exercise buys a document, not a capability. 21 capability-days out of 2,920, and the report was stale before it was circulated.
  • A drill in the calendar has already done the slow part of the incident. It measures 7.8% of the clock. Run two unannounced ones a year purely to learn the factor.
  • Rehearsing the repair is worth half a point; rehearsing the detection is worth thirty-five. Spend accordingly.
Term Meaning
Currency Whether a capability has been demonstrated recently enough for the demonstration still to be evidence.
Verified-day coverage The share of the year in which a capability is current — what an exercise programme actually buys.
Invalidating change A change to anything a capability depends on. Its rate λ is countable from git and the config history.
Discussion-based vs operations-based HSEEP’s own split. Only the second kind can make a capability current.
Announced-drill factor How much a scheduled rehearsal understates the real event. Measured, not assumed.
Gap A recovery path that is broken and untried. Gap-days, not gap count, is the exposure.
Discovery vs currency The two products an exercise sells; they want opposite cadences.
Steady-state hypothesis From the chaos principles: “focus on the measurable output of a system, rather than internal attributes.” For a data platform the measurable output is an answer.

14How these numbers were made

One deterministic generator (no RNG, no seed, no wall-clock input) over a synthetic year on a platform shaped like Lesson 30’s: 234 objects, 452 edges, 18 served answers of which 8 are tier-1, 11 external systems, 120 incidents.

Counted: every λ, every currency series, every coverage figure, the gap-days, the allocation table, the instrument’s distribution.

Carried in from earlier lessons, where they were measured: the four phase shares and the 189 h mean (L28); the 1.33 h alert and 16.5-day human detection lags and the 120-incident population (L25/L28); the 4 min 12 s restore and the 20.06 h / 69.95 h rebuild pair (L30); the €609 mutation-canary licence (L31).

Inputs: the change rates; €72/h loaded; build and upkeep hours for the automated drills; six unrehearsed repair durations; the 4% gap probability; the 21-day AAR lag. The repair durations are load-bearing, which is why the robustness table sets them to zero.

Checks. The closed form was validated against the simulation with phase-averaging (an earlier version compared a single annual draw against an expectation and reported up to 4 pp of pure sampling noise). The 0.00% was attacked as a round number before publication and survived with its expectation published beside it. The published SQL was re-run in DuckDB: 125 checks, 125 matched, 0 failed — one disagreement found on the way, Python’s sorted(x)[len(x)//2] against SQL’s quantile_cont(0.5) on an even-length list (the Lesson 18 median trap). dbt clone was actually run rather than described. Sources read at source: 14 CFR 61.57 (eCFR), the HSEEP doctrine 2020 revision, the chaos-engineering principles, the BigQuery table-clones documentation, and the clone macros from the shipped dbt-adapters 1.24.5 and dbt-bigquery 1.12.1 wheels. Round trip: both source files extracted from the rendered page, run in a clean directory, 95 top-level keys / 755 scalar values, 0 differences, figures.json md5-identical across runs and directories.

Two planned claims lost. The lesson was drafted to argue that unannounced drills should replace announced ones — costing the interruption showed they are a tax on the attention being measured, and the surviving recommendation is two a year used only to calibrate the rest. It was also drafted to argue for weighting the nightly rotation towards tier-1; that is right, but for the opposite reason to the one planned — the allocation with the worst mean coverage is the best one, and the metric ranks the three options exactly backwards.

No new failure shape. Like Lessons 25 and 31 this is a measurement of an instrument, not of the data. The scoring in § 33.10 is a judgement, laid out row by row so it can be argued with.

Artifact (visual version with six figures): published 2026-09-21.

Back to top