You Already Built the Tap

When I replaced a 150-endpoint MuleSoft implementation, the reverse proxy we put in front of it recorded every request it received and every response that came back. We needed those recordings for one job: diffing the old system against the new one, route by route, until the differences went quiet.

That's how I thought about the proxy at the time. It was scaffolding for the migration, and scaffolding comes down when the building is finished.

I've been rethinking that, because of a conversation that hasn't finished yet. It's the same one behind Native Means You Were Born There and its follow-up: a company that predates AI, trying to work out what it should actually do about it. The more time I spend on it, the more I think the proxy was the most useful thing we built. We just only used it for the migration.

A sharper diagnosis

In A Slogan Is Not a Diagnosis, I tried to name what's actually in the way. I still think that list is right. But it describes the symptoms. Underneath them is something more structural, and I can say it more precisely now:

The enterprise's legacy architecture and organizational structure are engineered for deterministic, tightly coupled, and siloed execution, which fundamentally chokes the high-velocity, probabilistic, and semantic data consumption required by AI.

That's a mouthful, so here's the plain version. These systems were built to do one thing correctly, the same way every time, for the one department that owns them. They are very good at that. What they were never built to do is let something else watch everything they do, all the time, and try to make sense of it. That's the thing AI actually needs from them.

Nobody designed that out on purpose. It just never came up when they were built.

Why getting the data out is the hard part

The textbook answer is change data capture: hook into the database's transaction log, and every change becomes a stream you can feed somewhere else.

With a decades-old system, that answer usually runs into one of three walls:

So the data stays where it is, and the AI conversation stalls on "we'd need to get at it first."

The tap was already there

Here's what I keep coming back to. If you've already put a proxy in front of the system, the way we did for the migration, every meaningful interaction with that system is already passing through something you control. Every request other systems make of it, and every answer it gives back.

That is a record of what the system actually does. Not what the documentation says it does, and not what the people who built it remember. What it does, on real traffic, today.

Repointing it for this purpose is a small change:

The property that matters most is the same one that made the migration safe: the legacy system never finds out. If the queue is down, the copy is dropped and the request still goes through exactly as it did yesterday. The live path doesn't get slower, doesn't get a new dependency, and doesn't get a new way to fail. You haven't touched the database, and you haven't asked the owning team to change a line of code.

It's change data capture by watching traffic, for systems that were never going to give you change data capture any other way.

What goes in the reservoir, and what comes out

Once the traffic is accumulating, two things become possible that weren't before.

The first is sense-making. Turn those recorded interactions into embeddings and they become searchable by meaning, not just by field name. Questions like "which customers keep hitting this error," or "what does a normal week of orders look like," stop being a report request that takes three weeks and become something you can just ask.

The second is that the system starts producing events it was never designed to produce. In Stop Polling, Start Listening, I wrote that for older systems with nothing to listen to, the discipline is turning "I saw a difference" into an event of your own. The tap is that, at the scale of the whole system. A response that confirms an order was created is an OrderWasCreated event, whether or not the system that sent it has ever heard of one. Publish those events where other systems can subscribe to them, and nobody has to ask the old system what happened anymore. They get told as it happens, AI included. The old system still only answers questions. The tap does the announcing for it.

What it can't see

I want to be careful here, because this is still an approach I'm working through, not something I've watched hold up in production. There are real limits, and some of them are sharp.

It only sees what crosses the wire. Nightly batch jobs, stored procedures that fire on their own, anyone editing rows directly in the database: none of that passes through the proxy. The reservoir is a record of the system's conversations, not of its internal state. That's a real gap, and pretending otherwise would give people a confident answer built on half the picture.

Traffic doesn't explain itself. I ran into a version of this on the SOAP migration: every operation arrives as the same HTTP POST to the same endpoint, and you have to open the envelope to know what was even being asked. A reservoir full of payloads you can't interpret is just a very expensive log. Somebody still has to do the unglamorous work of mapping what each interaction means. That's the real work it always was.

A recording is a liability, not just an asset. In a migration, the recordings have a job with an end date: prove the cutover. A reservoir is the opposite: its whole point is to keep things. That means customer data, possibly credentials in headers, all in a place that didn't exist before. Retention, access, and redaction have to be decided before the first copy lands, not after someone asks what's in there.

Fire and forget means some copies get lost. That's the price of never slowing the live path, and it's the right price. But it means the reservoir is for understanding the system, not for being the system of record. The moment anyone starts treating it as the source of truth, it quietly becomes a production dependency with none of the care a production dependency deserves.

Where this leaves me

I still don't have this resolved with the client. What's changed is that I no longer think the first step is buying something.

If the guiding policy from the last post was find out what's already true before deciding what to buy, this is the most literal version of it I've found. The proxy doesn't guess what the old system does. It watches, without ever getting in the way, and writes it down.

We built one of those to get off MuleSoft, and I thought of it as scaffolding the whole time. Next time, I'd build it as if it were staying.