FP&A The FP&A Department · 2 / 3

The craft: building models, and keeping them alive

Bookshops overflow with advice on how to build a forecast. Almost nothing is written on how to keep one alive for three years, which is most of the job, and most of where it goes wrong. How planning models are born, why they decay, the one person who secretly understands them, and what AI changes about all of it. Part 2 of a short series on the FP&A department.

An intricate ink illustration of a planning model drawn as a patched, Frankenstein-like machine labelled THE MODEL, plastered with notes reading DEFINITION DRIFT, HARDCODED PATCH, 11PM SHORTCUT and FINAL_v7_REALLY_FINAL, with a confidence gauge reading 18% beside a reality gauge reading 86%. One analyst tends it while a glowing robotic arm marked AI GENERATED MODULE (SPEED) bolts on new parts; to the right, a calm figure attends to MAINTENANCE IS THE CRAFT and STEWARDSHIP, THE SCARCE SKILL.
Nobody designs the Model. It is promoted (from answer to artefact to institution) one urgent Thursday at a time.

At the end of the last post I mentioned a person every FP&A department has and almost none has named: the one who actually understands how the model works. Not the official owner, the real one. The person everyone visits when the bridge doesn't bridge.

This post is about the work that person does, because in thirty years of working with finance departments I have come to believe it is the least written-about craft in the profession. Bookshops and blogs overflow with material on how to build a forecast. I have yet to find a serious piece on how to keep one alive for three years. And yet keeping it alive is most of the job, and most of where things go wrong.

How models are actually born

Here is the honest life story of nearly every planning model I have encountered, and I have encountered them across different industries, different sizes, different states of repair.

Nobody designs them. Somebody needs one, usually on a Thursday, usually for one specific decision. An analyst builds a quick file. It answers the question well, which is precisely what dooms it: a model that answers one question well gets asked a second question, then a fifth, then becomes the file the monthly review runs on. Within a year it is load-bearing infrastructure. Within three it has a name people say with a certain tone (the Model, Big Bertha, Frankenstein) and a folder of predecessors with names like FINAL_v7_REALLY_FINAL.

The model was never designed. It was promoted, from answer to artefact to institution, and at no point along the way did anyone stop to give it the structural attention an institution deserves. This is not a failure of any individual. It is what happens when a tool built for a question is asked to become a system, one urgent Thursday at a time.

Why models decay

Models do not break. They decay, and the distinction matters, because breakage gets fixed and decay gets lived with.

The decay has predictable mechanisms. The business changes faster than the structure. A new product line, an acquisition, a reorganisation: each one is absorbed by the model the way a house absorbs a new family member: a wall here, an extension there, never a new foundation. Definitions drift. Gross margin in the sales tab quietly stops meaning what gross margin means in the consolidation tab, and nobody notices until a board meeting does. Shortcuts compound. Every hardcoded number, every "I'll document this later," every patch applied at 11pm before a deadline is a small loan taken out against the model's future. Decay is simply the interest coming due.

And then there is the quiet arithmetic of error itself, which deserves a section of its own, because the research here is both sobering and, in an important way, kind.

What the research says about errors

The most thorough body of work on spreadsheet quality comes from Raymond Panko and the researchers around the European Spreadsheet Risks Interest Group, accumulated over decades of field audits and experiments. The headline findings:

Field audits of real, operational spreadsheets (the ones companies actually run on) have found errors in the overwhelming majority of those examined; the more recent audits, using better inspection methods, found errors in roughly nine out of ten. Measured per cell, undetected error rates run at a few percent of formulas.

Before anyone reaches for blame: the same research is clear that this is not a finance problem or a competence problem. Humans performing any nontrivial cognitive work (writing, programming, calculating) produce undetected errors at a few percent per action. Software developers, with their compilers and test suites and code reviews, show comparable rates per line of code. A large model is simply thousands of nontrivial cognitive actions stacked on top of each other, and the arithmetic of being human does the rest.

The finding I find most instructive is about confidence. In one study, people who had just built a spreadsheet were asked to estimate the probability it contained an error. The average guess: 18%. The measured reality: 86%. That gap (between how sure we are and how right we are) is, in my observation, the single most expensive number in this entire series. The errors are survivable. The certainty that there are no errors is what costs real money, as a handful of famous cases have demonstrated: an academic paper that shaped austerity policy undone by a range that omitted five rows; a major bank's risk model, run on pasted-together spreadsheets, understating risk ahead of a multi-billion-dollar trading loss.

The professional conclusion

It is not "trust nothing." It is: review is not optional, lineage must be visible, and any process that relies on one person's certainty is a process waiting for its invoice.

The model owner: the institution of one

Which brings us back to the person from the org-chart footnote.

In nearly every department I have worked with, the planning model's true documentation is a human being. The official documentation, if it exists, describes what the model does. The person carries the why: why depreciation is handled in that odd way (an auditor insisted, in 2019), why the France numbers route through a side table (an ERP migration, half-finished), why you must never, under any circumstances, sort column D. None of this is written down, because it accreted in the same urgent Thursdays the model did.

This person is usually excellent, usually modest, and usually the largest unbudgeted risk in the finance function. When they are promoted, poached, or simply take three weeks of parental leave during planning season, the department discovers what it actually owned: not a model, but access to one person's memory of it. I have watched the handover attempt many times. It follows a pattern: two weeks of knowledge-transfer sessions, a hopeful wiki page, a successor who six months later has rebuilt a quarter of the model from scratch, not out of arrogance, but because rebuilding what you understand is genuinely easier than inheriting what you don't.

Here is the sentence I would ask every finance leader to sit with:

A model only one person understands is not owned by the department. It is on loan from that person, and the loan can be recalled with four weeks' notice.

Maintenance is the craft

So what does keeping a model alive actually involve? Less glamour than building, and more discipline. From the departments that do it well (they exist, and they are recognisable within an hour of walking in) a few practices recur:

Definitions live in one place and are owned. Gross margin means one thing, defined once, referenced everywhere. The moment a definition can be locally improvised, drift begins. The departments that get this right treat their definitions the way accounting treats the chart of accounts: as governed vocabulary, not personal preference.

Lineage is inspectable. For any number in the output, someone other than the model owner can trace where it came from and what touches it. If tracing a number requires the one person, see above.

Assumptions are written where they live. Not in a separate document nobody updates, next to the thing they explain. The why must travel with the what, or it travels out the door with its keeper.

The model gets reviewed the way code gets reviewed. Periodically, by someone who didn't build it, with the explicit licence to ask naive questions. The research on inspection is humbling. Even careful reviewers catch perhaps half of what's there, which is an argument for doing it regularly, not for skipping it.

Changes have a record. What changed, when, by whom, and why. Version seven of FINAL_REALLY_FINAL is not a change record. It is an archaeology site.

None of this is exotic. All of it is unglamorous. Most of it is absent, not because finance people are undisciplined, but because every one of these practices costs time now to save time later, and the Tuesday afternoon from Part 1 has no time now. Maintenance is the part of the craft that organisational gravity works against, which is exactly why it distinguishes the departments that have it.

And AI?

The closing question of every post in this series, from this post's angle. And this is the angle where the answer has the most edge.

AI makes building dramatically faster. It does not make owning anything faster. A capable assistant can now draft in an afternoon what took an analyst two weeks. That is real, and it is welcome. But notice what it does to the economics of the craft: when construction becomes cheap, the bottleneck moves. It moves to exactly the things this post has been about: understanding, lineage, definitions, stewardship. A model you generated before lunch and cannot explain by dinner has not solved your problem; it has automated the creation of the problem.

The decay mechanisms do not care who built the model. Definitions drift in generated models too. The business still changes faster than the structure. And the overconfidence gap (18% versus 86%) has, if anything, a new variant: confidence in output one did not personally construct is even harder to calibrate than confidence in one's own work.

So my honest expectation, watching this unfold across the departments I work with: AI will divide FP&A teams not into those who use it and those who don't (everyone will use it) but into those who let it accelerate building while they keep ownership of understanding, and those who quietly trade understanding away for speed. The first group gets the compounding benefit. The second group gets a faster route to the institution of one, except this time the institution is a system nobody can interrogate at all.

The new scarce skill is not prompting. It is stewardship: the ability to keep a model understood, governed, and alive while it is being changed faster than ever. The departments that already had the maintenance craft will find AI makes them formidable. The departments that didn't will find AI makes them faster at accumulating debt.

What remains, once the building accelerates and the stewardship holds, is the part of the job no tool has ever done: standing in front of the people who decide, and turning the model into a story they can act on. That is Part 3.

Part 3, "From numbers to narrative: storytelling as a core function", is now live.

Sources referenced
  • Raymond Panko, What We Know About Spreadsheet Errors and related field-audit research: error rates in operational spreadsheets, per-cell error rates, and the overconfidence findings (arxiv.org/abs/0802.3457)
  • European Spreadsheet Risks Interest Group (EuSpRIG): error-detection rates and spreadsheet risk research (eusprig.org)
  • Reinhart & Rogoff, Growth in a Time of Debt spreadsheet error (Herndon, Ash & Pollin, 2013); JPMorgan Task Force Report on the 2012 CIO losses: documented cases of consequential model errors

A model finance owns, and a machine can read

Novi is built around the things this post is about: definitions in one place, lineage you can inspect, assumptions that travel with the numbers, and an AI that reads the whole model instead of chatting on top of it.