
Three Records
enterprise data governancedigital ecosystem architecturetechnical debt managemententerprise architecturesystem integration challengesdata silosoperational transformationbusiness process governancedigital operating modeltechnology governancetransformation leadershipIt was late afternoon but we had decided a pilsner of beer was earned after the efforts of the day. I sipped as I listened and commiserated with my Client.
"Three records!" he spat out, for the third time that hour. "That customer existed in three different records! And one completely stale! And that's why we charged him the wrong thing!"
---
The day had started with him calling me at about half past eight, which was already a signal, because he is not a man who calls before nine unless something has gone properly wrong.
A customer had cancelled a pre-paid service package. He'd sold the car - moved to another country, I think, or maybe the car was traded in, I don't remember now and it doesn't matter. The package had a couple of services left on it, and the terms allowed for a pro-rata refund on whatever hadn't been used. Straightforward enough. Except the refund that came out was noticeably less than it should have been, the customer had done his own arithmetic, and he was not shy about it.
"Your system charged him wrong," is roughly how it was put to me. Not unkindly, but not warmly either.
I was on my way to another appointment when I took the call, and I cancelled it before I'd finished the conversation. Not because I thought it was my system - although at that point I genuinely didn't know - but because of the shape of the error. A pro-rata refund is arithmetic. It's the least clever thing the platform does. If the arithmetic was wrong, then the arithmetic wasn't the problem, and whatever was actually wrong had been sitting there quietly for however long, only surfacing now because somebody happened to cancel a package early, which almost nobody ever does.
That's the thing about fringe cases. They aren't rare because they're difficult. They're rare because the conditions rarely line up. And when the conditions do line up, they don't fail gently.
The walkthrough
He met me downstairs and we went through it together. I want to be fair to him here - he was not pleased, and he had every right not to be, but he still walked me through it properly rather than just handing me a printout and folding his arms. We have that kind of relationship. It has taken years.
The record was there. The consumed services were listed. The maths, given those inputs, was correct.
Which meant the inputs were wrong.
So I opened my laptop and asked the only question that was left: where did this record come from?
Now, some background on how the thing is wired. During discovery - the long one, the one I've written about before, the one that took longer than the build and made everybody nervous - it was established that all transactional records would come from System 1. That was not my decision, it was theirs, and it was the right one because System 1 was the transactional system. It held the money.
There was also a System 2. Older, mandated by their global principal, and generally regarded as a place where things went to be recorded rather than to be used. It was strictly specified - in writing, in a document with signatures on it - that System 2 held no transactional data. My platform therefore read from it for other purposes and ignored anything that looked like money.
Except somebody, about eighteen months earlier, in a different department, had started using a module in System 2 to keep track of a particular add-on service. Not maliciously. Not stupidly, either - they had a genuine gap. They needed to track that item, they had access to that module, nothing else existed for it, and so they used what was in front of them. If you have ever worked in operations you have done exactly this, and so have I.
What they did not know - could not have known because nobody had told them (because why would anyone tell them?) - was that the item they were now recording in System 2 was already being captured in System 1, amalgamated into a line item under a different name. And what they really could not have known was that my platform was reading that module.
So the item was counted twice. In the ordinary run of things this was invisible. It affected nothing anybody looked at. It only mattered if you ever needed to total up what a customer had consumed and subtract it from what they'd paid - which is to say, it only mattered on a cancellation, which is the one thing that almost never happens.
I did not have to explain this - both the Ops Director and myself had come to the realisation together and were looking at each other in disbelief. Behind us, there was a silence of the kind where somebody is doing organisational arithmetic in their head about which department this belongs to.
Then it got worse
The obvious fix was sitting right there. Either my platform stops reading that module, or they stop using it. The first was ten minutes of work. The second was a bad idea, because the gap they'd been filling was a real gap and taking away their workaround without replacing it just moves the problem somewhere I can't see it.
So: ten minutes of work. Fix it, apologise to the customer, go home.
Except the numbers still nagged at me.
I did the calculation by hand - actually by hand, on paper, because at that point I temporarily didn't trust anything with a screen. Removed the duplicate. Ran it again. It was closer.
But it was not right.
I was willing to believe maybe my manual arithmetic skills were rusty, so decided for good measure to run it again on an Excel sheet. Same.
There was still a gap. Small, and consistent enough that it was obviously a rate and not an error, and once you're looking for a rate it doesn't take long. One of the service prices flowing into System 1 was arriving from a third party's system entirely - a vendor integration nobody in the room had thought about in years - and it was arriving with tax applied at six percent.
Six percent. GST. A tax that had not existed in this country for the better part of a decade.
Not the SST that replaced it. Not the rate SST later became. A ghost rate, from a regime that was abolished, sitting inside a live price feed, quietly making every single one of those line items slightly wrong for years - and invisible, because who audits a price that looks approximately correct?
We sat with that for a moment.
What I didn't do
I did not fix it that day.
My team could have - it would not have been hard, and would likely have taken less than half an hour. But I had just found two independent faults in a system I thought I understood, in the space of four hours, and both of them had been sitting there for years without anybody noticing. I felt that the right response to that is not to fix the two we found. It is to assume there is a third that I might not know about at all.
So I told him: tomorrow, we audit this together, properly, both of us in the same room going through it line by line, and in the meantime I'll work on what the flow should look like rather than patching what it currently is. He agreed - I think with some relief, actually, because it meant he wasn't going to have to explain a same-day fix to anybody upstairs and then explain a second one next month.
It had been a long, unglamorous, fairly demoralising day for both of us, and rather more demoralising for him than me, because somewhere in the middle of it he had worked out that most of this originated inside his own organisation and not inside my software.
I was packing up when he said he needed a drink and asked if I wanted one.
I knew that it was his way of apologising for his assumption this morning.
The part that isn't about the software
I did the long discovery on that account. I walked the floors, I sat with the operations people, I asked the questions. And I still didn't catch this, because the workaround that caused it was introduced eighteen months after I'd finished, by a team I had no reason to speak to, solving a problem nobody had escalated, using a module that a signed document said contained nothing of interest.
I'm a vendor. I have no seat in that organisation, and I cannot enforce an SOP there. I don't attend their internal meetings, and nobody is obliged to tell me when a department invents a new use for an old system. Any ecosystem architect worth their salt might see a great deal - but only ever what has been made known to them, and only as of the day they were told.
And this is not really a story about one client. Most companies run this way, and Groups especially so, particularly the ones in the middle of consolidating. Systems get bought to do one thing and quietly asked to do four. A department solves its own problem in an afternoon - increasingly, these days, with AI, and rather impressively - and has no way of knowing that three floors away, something is reading the table they just started writing to.
So what's the actual governance answer? It isn't nothing changes without IT approval. That fails immediately, partly because it strangles exactly the initiative that every leadership team is currently telling its people to show, and partly because a lot of IT departments are not equipped for it anyway - they run the infrastructure, they don't hold the commercial logic or the operational reality or the data governance picture in one head at the same time.
Which points at a role that, in most organisations, simply does not exist. Somebody whose actual job is the shape of the whole thing - technical enough to read the wiring, commercial enough to know what a wrong line item costs, operational enough to know why the workaround was invented in the first place, and senior enough that when they walk into a department and ask what that module is being used for, they get an answer.
Most companies don't have that person. They have some of it distributed across four people who each see a quarter of the picture and meet quarterly.
Technical debt, and what's coming
Somebody I know had their expense claim paid into a stranger's bank account recently. Same first name, different person, different account number - and their salary, meanwhile, went to the right place, which is somehow the more alarming detail, because it means the organisation held two separate versions of them and never noticed the versions disagreed.
One person. One account. That should be the whole of it. Whether the money is a salary or a claim, whether it comes through HR or through an approval chain, both paths should converge on the same record of the same human being before anything moves. When they don't, you don't have a payments problem. You have two people in your database wearing the same name.
I could cite half a dozen more of these, from clients, from my own experience, from friends in finance functions who will tell you about it at length if you let them.
Technical debt is the term that often gets thrown around. But much of this isn't technical debt at all. It's organisational debt that happens to have become encoded in technology. But I think what we're actually going to see, over the next few years, as everybody sprints to bolt AI and automation onto estates that were already held together with goodwill and Excel, is a great deal more of this - and faster, and at greater volume, and discovered later.
Not because the tools are bad. Because we keep pointing extraordinary tools at a foundation nobody has audited in years, and then acting surprised at the output.
The customer got his refund. Correctly, in the end.
---
The scenario in this article is fictional - a composite drawn from real incidents and several different people. The tax rate, unfortunately, is not.
I write these as they happen - on discovery, documentation, AI and the shape of systems. New pieces go out on LinkedIn first.
Follow on LinkedIn RSS