Roadmap To Be A Data Engineer / Lesson 28

Lesson 28 Ownership Fundamentals §8 Governance / §7 Serving (+§5, §6) About 20 min read

Who Owns This Table?

A seven-hour fix takes twenty days because nobody owns the table.

The alarm fired in 23 minutes. The fix was 7.3 hours of engineering. The company was wrong for 19.83 days. Nothing in between was an engineering failure.


1The scene

Saturday 4 July, 15:30: an app release goes out. At 15:53 — 23 minutes later — a value-domain test on int_listing_state fails. The listing service has begun writing relist where the previous build wrote create, so every republished item looks brand new: item age resets to zero, the markdown queue empties, category mix shifts.

This is exactly the instrument Lesson 25 argued for, and it works perfectly. The fix is four lines in the mobile client plus a coalesce in a staging model — 3.6 h of Product Engineering’s time and 3.7 h of Analytics Engineering’s, split between two people who both know immediately what to do.

It ships on Friday 24 July at 11:42.

release → alarm 23 minutes
engineering in the fix 7.3 h
alarm → shipped 19.83 days
elapsed ÷ work 65.2× (work = 1.53% of the clock)
served answers wrong 17, of which 8 are tier 1
damage 8,095 served-object-hours

Where the 19.83 days went: 94.1 h finding who should look at it, 284.0 h waiting for a team to pick it up, 23.3 h of working window (of which 7.3 h is typing; the rest is nights and a weekend), 74.4 h waiting for a release train.

On the Monday, the CTO asked in the incident channel: who owns int_listing_state? Five people answered, in good faith, and gave five different answers.

2Five certificates, one table

Record Answer Detail The catch
dbt group Zeynep Winter group listing_domain still here, and has never committed to it
Runbook page Nele Vogel last edited 2025-06-10 left the company
git history Tobias Demir 45% of commits, 4 contributors never named as owner anywhere
On-call rota svc-data-warehouse routes to Analytics Engineering, 3 people a rota is not a person
The reader Growth Analytics the team on the dashboard where the number is wrong owns no model on this path

Four of the five point at Analytics Engineering, so the team was never really in dispute. It did not help: the change that had to be made was in the mobile client, owned by Product Engineering, which none of the five records mentions at all.

Across the whole warehouse:

Teams named by the five records Tables Share
2 26 36.1%
3 34 47.2%
4 10 13.9%
5 2 2.8%
any single team 0 0.0%

Zero is the kind of number that is usually a bug. It is not one here, but it is nearly structural: two of the five records are fixed by something other than the table (the pager rota by which service the table sits in; the reader’s belief by whichever dashboard they had open). Only 2 of the 72 tables are even capable of unanimity, and neither achieves it. The honest claim is not “nobody agrees” — it is that agreement is not something these records are built to produce.

3What the tool actually records

Read from the shipped source of dbt-core 1.12.4, not the documentation:

Where What the source says What that means
model.py Model has no owner field a model cannot have an owner; only a group can
components.py ParsedResource.group: Optional[str] = None and the group is optional, defaulting to none
owner.py Owner(email: Union[str, List[str], None] = None, name: Optional[str] = None) an owner is two optional strings; a list of addresses is explicitly allowed
unparsed.py "Group owner must have at least one of 'name' or 'email'." a bare group alias is a legal owner, and nothing checks it reaches a person
exposure.py Exposure.owner: Owner — required, no default a dashboard, notebook, analysis, ML model or application must have an owner
components.py Contract.enforced: bool = False contracts are off unless somebody turns them on
model.py ModelConfig.access = AccessType.Protected every model is visible to its own project by default

dbt requires an owner for an answer and provides no owner field at all for a table. Ownership of an exposure is mandatory; ownership of a model is optional and indirect. Almost nobody uses it that way round.

In this warehouse:

Count Of
Tables 72 —
… dbt models 60 72
Models with no group at all 17 60
Models whose group owner is a mailing list 21 60
… all the same list, 6 members 21 60
Models whose group owner is a named person 22 60
… who has left 2 22
… who has never committed to the model they own 13 22
Models with contract: {enforced: true} 2 60
Tables with no runbook 36 72
Runbooks written by a leaver 5 36

4A table is not an answer

Suppose you fixed all of that tomorrow — one live, named human against every table. The 18 things people actually read stand on:

  • minimum 4 teams (3 internal + a vendor)
  • median 5
  • maximum 7
  • answers ownable end to end by one team: 0

There is no assignment of tables to owners that gives any of these answers a single owner. Ownership of a table and ownership of an answer are different objects, and the second cannot be derived from the first. Of the 18 served objects, 6 are declared in the dbt project as exposures (and therefore have an owner, because dbt insists); the other 12 exist only in the BI tool, where “owner” is whoever clicked New Dashboard.

A table owner answers “is this model correct and maintained?” An answer owner answers “is this number right, and if not, who is making it right today?” The second is a job with a queue and a budget attached. The first is a line in a YAML file.

5Where the time actually goes

120 incidents over twelve months, every hour attributed:

Phase Hours Share What is happening
Finding the owner 6,199 27.3% asking, being told “not us”, asking again
Waiting to be picked up 11,584 51.1% sitting in a backlog until the next sprint boundary
The working window 1,768 7.8% of which 712 h is typing — 3.14% of the total
Waiting for the release train 3,129 13.8% correct code, merged, not in production

712 engineering hours against 22,680 elapsed hours — a ratio of 31.9×. Even inside the working window, only 40.3% of the clock is someone working.

Teams the fix crossed n Median days Work each
1 36 3.08 3.0 h
2 72 6.94 6.4 h
3 12 16.58 11.7 h

The work grows 3.9×; the clock grows 5.4×. The extra time is not extra difficulty.

Two honest limits. The engineering hours were drawn independently of how many teams a fix crosses, so the absence of that relationship is built in, not discovered. What is not built in is the size of the gap, which falls out of four published facts: a fortnightly sprint, a weekly release train, a monthly close and an eight-hour day. And the robustness check: make every fix ten times harder and the work still reaches only 18.4% of the elapsed time; a perfect organisation (instant routing, instant acceptance, only the release trains left) still runs at 8.1×, median 1.15 days.

6The calendar decides the fix date

The same defect and the same fix, replayed from each of 84 consecutive detection dates:

  • range 9.83 – 22.83 days (2.32×)
  • found 14 July: 9.83 days. Found 15 July — one day later: 22.83 days. A gap of 13 days, recurring every fortnight, because the routing delay landed a few hours after sprint planning instead of a few hours before.

And the sharper reading: the real incident was found on 4 July and shipped on 24 July. Had it been found on 14 July — ten days of wrong answers later — it would have shipped on exactly the same day, at exactly the same hour. For those ten days, detection bought nothing.

Lesson 27 measured a queue made of slots. A sprint is a scheduler with a fortnightly quantum and no preemption, and the thing queued is a repair rather than a query. Every property holds: the median is fine, the tail is terrible, the variance is invisible from inside the job.

7Detection was never the bottleneck

They took Lesson 25’s advice. Four structural checks went live on 1 April, and worked: median time to detection for the newly covered classes fell from 123.8 h to 0.5 h — 247.6×.

Counterfactual: give every incident in the year the detection latency of an instrumented one (median MTTD 56.8 h → 44 minutes, 77.5×), change nothing else, re-run the queue.

  • Wrong answers fall by 24.6%. That is the whole ceiling on instrumentation.
  • The share of damage happening before anyone knew goes 38.0% → 0.57%, leaving 99.4% of the damage sitting in a queue.

Lesson 25 measured 88.59% of its damage pre-detection and concluded, correctly, that the company had to look harder. Look hard enough and the number inverts. The second half of this problem is not an observability problem at all.

8Every boundary in the warehouse

A handoff is an edge in the DAG whose two ends belong to different teams. 54 of 100 dependencies (54%) cross a team boundary. 2 have a contract.

Team Depended on by Depends on
Analytics Engineering 27 11
Data Platform 12 12
Product Engineering 9 0
External vendors 3 0
Growth Analytics 2 8
Finance 1 7
BI & Reporting 0 16

A contract on a source is written by the team that emits the data and protects the team that reads it. Product Engineering has 9 outbound boundary edges and 0 inbound: every contract in this warehouse costs them something and protects them from nothing. That is not a culture problem, it is an accounting one.

Of 120 incidents, 84 (70%) needed more than one team and carried 75.8% of the year’s wrong answers. A further 24.2% originated in a table the dbt project does not describe at all — a source, where there is no owner field to fill in because there is no model.

9Who breaks it, who pays for it

Team Share of damage Share of repair work People
Product Engineering 41.1% 21.3% 11
Analytics Engineering 34.4% 37.8% 4
Data Platform 14.1% 23.9% 5
Growth Analytics 4.3% 5.8% 3
External vendors 3.5% 0.0% —
BI & Reporting 1.9% 6.3% 3
Finance 0.7% 4.9% 2

Of the six internal teams, exactly one exports more wrongness than it repairs. This is the structural reason data quality feels like a data-team problem to everyone except the data team — and why “shift left” is a sentence rather than a plan.

10Six interventions, each run through the same year

Change Wrong answers left Reduction Median repair
L0 Nothing (today) 311,754 — 6.32 d
L1 An owner on every model 265,215 −14.9% 5.09 d
L2 Every team takes Sev-1 and Sev-2 interrupts 279,883 −10.2% 4.25 d
L3 One data team (Platform + Analytics Eng merge) 277,789 −10.9% 5.19 d
L4 Contracts on the source boundary 179,702 −42.4% 5.37 d
L5 A named owner for each tier-1 answer 176,720 −43.3% 1.84 d
L6 All of the above 92,276 −70.4% 0.79 d

The cheap fix is worth real money and is not the answer. A live, named owner on every model buys 14.9% — about what you would expect, given that finding the owner is 27.3% of the elapsed time. First thing to do; last thing that finishes the job.

The metric everybody reports ranks the options backwards. By median repair the order is L2 < L1 < L3 < L4; by damage removed it is very nearly the reverse (Spearman ρ = −0.80). Taking interrupts has the best median (4.25 d) and the worst damage reduction (10.2%); contracts have the worst median (5.37 d) and the best damage reduction (42.4%). The reason is concentration: 98 of 120 incidents touch a tier-1 answer and carry 96.6% of the damage, and a median is precisely the statistic that cannot see them.

The narrow fix beats the broad one. A universal Sev-1/Sev-2 interrupt obligation buys 10.2%. Eight named people — one per tier-1 answer — with the authority to interrupt anybody on behalf of their answer buys 43.3% and takes the median from 6.32 days to 1.84. Same mechanism, aimed.

What the interrupts cost, priced:

Interruptions a year Engineering hours Share of capacity
Product Eng 8 57.2 0.32%
BI & Reporting 9 27.8 0.58%
Growth 7 27.1 0.56%
Finance 8 25.3 0.79%
All teams 32 137.3 0.31%

137 engineering hours a year — 0.31% of six teams’ capacity — buying 232 object-hours of correct answers per hour spent. That trade has never been written down anywhere, because no forum exists that owns both sides of it. (The answer-owner rung costs more and buys more: 55 interruptions, 227 hours, 0.51% of capacity.)

And the price of the contracts row, stated: a contract does not make a bad release go away, it makes the load fail instead of succeeding quietly. In this year that converts 42.4% of wrong answers into 113,884 object-hours of stale ones — a good trade, since stale is a state every monitor already sees and nobody acts on wrongly, but a trade. The stale figure is an upper bound (a blocked pipeline gets escalated in a way this simulation does not model).

11What to ask on Monday

  1. “Name the eight numbers we would stop the week for, and the person who owns each one — not the table, the number.” If the second half of each pair is a team, a rota or a mailing list, you have table owners and no answer owners.
  2. “For the last five incidents: what time did we know, what time did someone start, what time did it ship?” Three timestamps, five tickets. The gap between the first two is your routing and queueing cost.
  3. “Which of our sources can change shape without our build failing?” Here it was 52 of 54. Not a culture question — a list.
  4. “When a fix needs Product Engineering, what happens, concretely, this week?” You are listening for a sprint boundary and a release train. Those two facts set the floor on every cross-team repair, and nobody chose them with data in mind.
  5. “Who decides what ‘active seller’ means, and where is that written as data?” Lesson 20’s question, asked as an ownership question.

12Hands-on (20 minutes, your own project)

1. Census your own paperwork. Run dbt parse, then:

jq '{
  models:            [.nodes[]    | select(.resource_type=="model")]                        | length,
  without_a_group:   [.nodes[]    | select(.resource_type=="model" and .group==null)]       | length,
  contract_enforced: [.nodes[]    | select(.resource_type=="model" and .contract.enforced)] | length,
  groups:            (.groups     | length),
  groups_with_no_human_name:    [.groups[]    | select(.owner.name==null)] | length,
  exposures:         (.exposures  | length),
  exposures_with_no_human_name: [.exposures[] | select(.owner.name==null)] | length
}' target/manifest.json

Then confirm the two facts §3 rests on, directly in the artefact:

jq -r '.nodes["model.YOUR_PROJECT.SOME_MODEL"]
       | "has owner key: \(has("owner"))   group: \(.group)   " +
         "contract.enforced: \(.contract.enforced)   access: \(.access)"' \
   target/manifest.json
# -> has owner key: false   group: null   contract.enforced: false   access: protected

2. Pick your three most-read numbers and list every team whose work has to be right for each to be right, vendors included. If any comes back with one team, you have found either a very well-organised company or an incomplete list.

3. Take the last five incidents and record four timestamps: defect entered, human knew, someone started, it shipped. Then compute engineering hours ÷ elapsed hours. Anything under 10% means here what it means there.

4. Reproduce the lesson: the generator runs on stock Python 3 with no packages. Change Product Engineering’s sprint from two weeks to one and re-run to see what a calendar change is worth.

13The eighteenth failure shape: the unclaimed

Event (06–10) / constant (11) / drift (12) / unwatched (13) / reversible (14) / unreproduced (15) / bundled (16) / transient (17) / extremal (18) / referential (19) / plural (20) / remote (21) / absent (22) / permitted (23) / detached (24) / unsettled (26) / contended (27) — and now the unclaimed.

The defect is found, named, reproducible and cheap to fix, and it persists because the repair is queued in a human system nobody is accountable for. Its signature: all of the damage is post-detection, which inverts everything the first seventeen shapes were about; the unit of failure is a handoff rather than a row, a job or an interleaving; and the difficulty of the fix predicts nothing while the calendar predicts almost everything. It is the only shape whose fix is not code.

14Takeaway and vocabulary

An owner is a capacity commitment, not a string. A name in a YAML file that carries no obligation to drop something is decoration. The eight answer owners are worth 43% precisely because the role comes with the right to interrupt — 55 interruptions and 227 engineering hours a year, 0.51% of capacity, priced and agreed in advance.

Own answers, not tables. No assignment of tables to people can give an answer a single owner: the tightest of the 18 things people read crosses 4 teams. dbt — which requires an owner for an exposure and offers none for a model — has been saying so all along.

A contract is what makes a boundary crossable without a meeting. 54 of 100 dependencies cross a team; 2 are written down. The other 52 are re-negotiated, informally and slowly, every time something changes.

Term Meaning
Data product An answer with a name, an owner, a published definition, an SLO and a contract with each of its inputs.
Data contract An enforced, versioned agreement at a boundary between two teams (dbt: contract: {enforced: true}, off by default).
dbt group The only way a model gets an owner, and it is optional (group: Optional[str] = None).
Exposure dbt’s model of a served answer — dashboard, notebook, analysis, ML, application. Owner mandatory.
Federated ownership Domains own their own data products under central standards; the platform team builds the road, not the cars.
Conway’s law Melvin Conway, Datamation, April 1968: “organizations which design systems … are constrained to produce designs which are copies of the communication structures of these organizations.”
Interrupt budget The share of a team’s capacity reserved, in advance, for work that arrives unplanned.
Handoff The unit this lesson counts: one transfer of a repair from one planning group to another.

15How the numbers were made

Assumed: the org. Six teams, 28 people, a 72-table warehouse with 100 dependencies, 120 incidents over the twelve months to 31 August 2026, and the five ownership records. Engineering hours are drawn independently of how many teams a fix crosses.

Measured: everything downstream. The repair queue is a discrete-event simulation whose only timing inputs are four things any company can read off its own calendar:

Team People Planning Takes interrupts at Ships
Product Engineering 11 2-week sprint Sev-1 only Thursday release train
Data Platform 5 1-week sprint Sev-1 and Sev-2 10:00 daily
Analytics Engineering 4 continuous any severity on merge
BI & Reporting 3 2-week sprint (out of phase) Sev-1 only on merge
Growth Analytics 3 2-week sprint never on merge
Finance 2 weekly council, not during month-end close never on merge

Finance is the exception worth naming: a metric definition that needs Finance and arrives on the 2nd waits longer than one that arrives on the 20th — a rule nobody wrote for data and everybody inherited.

Determinism. No RNG and no seed: every draw is a splitmix64 finalizer over string keys. Four primary sources were read rather than recalled — dbt-core 1.12.4 from the shipped wheel, the harmonised Community codes in Annex I of Directive 1999/37/EC (C.1 holder / C.2 owner / C.4(c) “is not identified by the registration certificate as being the vehicle owner”), Conway’s 1968 Datamation paper, and a real dbt parse on a small project to confirm the manifest query returns what it claims.

Checks. The generator asserts that the four phases of every incident sum exactly to its repair time, that the ownership counts partition the 60 models, and that no answer stands on fewer than two teams. The ten published queries were re-run in DuckDB over the generator’s own tables (72 tables, 100 edges, 120 incidents, 218 handoff steps): 71 checks, 71 matched, 0 failed. The generator was then lifted back out of the rendered page, run in a clean directory and diffed key by key: 1,136 values, 0 differences, byte-identical to source.

Sources. dbt-core 1.12.4 (dbt/artifacts/resources/v1/{model,components,owner,exposure}.py, dbt/contracts/graph/unparsed.py); Directive 1999/37/EC, Annex I; Melvin Conway, “How Do Committees Invent?”, Datamation, April 1968.

Back to top