Nobody owns whether the model is right
I wrote a couple of weeks ago about models that fail while every dashboard stays green. Someone on r/MLOps replied with a sentence that reframed it:
Your dashboards look fine because they were built to answer "is the service up", and nobody owns "is the model right".
He's right, and I had been writing about the wrong thing. Instrumentation isn't the bottleneck.
"Is the service up" has a pager, a threshold, a runbook and a person who gets woken up. "Is the model right" has none of those, so it gets answered eventually by a customer, or by someone who happens to look at a few outputs and thinks huh.
The usual state isn't that nobody owns it. It's that two people own half each. Platform knows the containers are healthy and the devices are online, and has no idea what a correct detection looks like on device 47 at four in the afternoon. The ML team can tell you in ten seconds whether a sample is wrong, isn't on the pager, and in a lot of places finds out about a rollout after it happened. Both assume the other has it.
Adding a fifth drift metric doesn't help. It just means that after the incident you can prove the signal was sitting there for eleven days.
The harder part is timing. The trustworthy answer to "is the model right" is a joined label or a business number, and it arrives on its own schedule rather than your rollout's. A bad model lands on 300 devices Tuesday, the number moves the following Monday, and that's six days of wrong output before your best signal notices. The authoritative metric is the right thing to own, and on its own it's too late to prevent anything.
Two layers work better. A weekly scorecard on the joined label, named owner, fixed cadence, reviewed whether or not anything looks wrong. That's ground truth and it catches slow degradation nothing else will. Then two or three cheap counters computed on the device: empty-output rate, per-class rate, the shape of the confidence histogram rather than its mean. Those are approximate and they'll occasionally cry wolf. Their job isn't accuracy, it's being fast enough to trigger a rollback before Monday. The tripwire gates the rollout, the scorecard grades it.
One thing about ownership. If the answer to who owns this is "the team," it's unowned. Teams don't get paged. It needs a name, a threshold agreed before the incident rather than after, and an action that's usually roll back first and investigate second.
All of that costs engineering time except the part that matters most. Pick the metric that best approximates "is this model doing its job," even a rough one, and put a person's name next to it. You'll improve the metric later. An imperfect metric with an owner beats a perfect one without.
Framing from u/dthompson_arch on r/MLOps, who also supplied the fix in five words: one engineer owns both.