TL;DR: Your MLOps model registry reflects voluntary check-ins, not the real production estate. The fix is not a stricter registry. It is an inventory built from inference traffic, billing, and code archaeology. The four-layer audit and 90-day consolidation playbook below turn 200 unknown models into a known, owned, and governed estate.
Key Takeaways: - The registry shows what teams chose to register, not what is running. The gap between registered artifacts and deployed models grows with every notebook export and untracked CI deployment. - Jupyter notebooks become the de facto production layer because they remove every friction point the registry adds. - An artifact-centric audit beats any off-the-shelf scanner, because it never assumes the registry is the source of truth.
The Registry Is a Lie, and Auditors Will Prove It

Your MLOps platform shows six registered models. Your engineering org has quietly deployed more than 200 in the last eighteen months, and almost none of them passed through the registry. This is MLOps model sprawl, and your CTO dashboard is lying to you.
The number on the registry is not a count of models. It is a count of voluntary check-ins. Every line in your registry represents a moment when a human or CI pipeline decided the model was important enough to record. Everything else is invisible.
Here is the typical gap: 6 registered, 200 in production. The ratio is not a coincidence. The registry's friction (CI gates, metadata forms, approval reviews) filters out the work that engineers consider routine. The "boring" fraud-scoring model, the "quick" recommendation tweak, the notebook export that solved a P1 incident at 2 a.m. None of these touched the registry. All of them serve real users today.
The risk compounds. An undocumented model has no lineage, no owner, no rollback path, and no performance baseline. When it drifts, no one notices. When it serves a biased output, no one logs it. We have written before about the governance gap boards miss when the inventory itself is wrong.
When regulators, customers, or your own risk team asks for an AI model inventory, the registry answer will be wrong. Not slightly wrong, structurally wrong. Six entries, where 200 are live. That gap is the story auditors will tell about your program.
But how did the registry get so out of sync? Nobody intended to bypass it. The drift happened inside Jupyter, cell by cell.
How Notebooks Quietly Became the Production Layer
The notebook was designed for exploration. In practice, it has become the production layer for a large share of enterprise ML. The shift is not malicious. It is ergonomic.
A notebook removes every gate the registry enforces. No CI run. No metadata schema. No approval review. The author runs a cell, sees a result, exports a checkpoint, and moves on. The friction is zero. The iteration speed is unbeatable. The model that solves a real problem ships before anyone asks whether it should.
The problem is that notebooks are not deterministic. The verified research on notebook-to-production pipelines is explicit. Cells execute sequentially in the UI, but nothing guarantees they always run in the same order. A data scientist might run cell 5 before cell 4, modify state midway, and forget the change. After 200 cells developed over several weeks, the author often cannot reproduce their own work from start to finish. Production needs the opposite: a deterministic, repeatable process.
This is what MLOps shadow IT looks like. It is not unauthorized SaaS or rogue cloud accounts. It is authorized Jupyter instances, fully inside the firewall, wired to live data, whose outputs are copy-pasted into serving code.
The notebook is already inside your perimeter. It already has credentials. It already runs. That makes this form of shadow IT more dangerous than the textbook kind. A traditional shadow asset lives outside the registry but also outside the data path. A shadow model is already serving production traffic, often with the same data rights as your registered flagship. The exposure is not theoretical. The audit trail goes dark the moment a notebook is exported, and the gap compounds from there.
If the problem is notebooks, the registry should catch them on the way out. So why does every platform miss the same models?
Why Registries Only See What Teams Choose to Register
A model registry does exactly what its name says. It tracks registered artifacts: model versions, metadata, lineage, and ownership. The verified MLOps reference confirms this scope. Every field exists, every API works, every schema is enforced.
The registry is not broken. The assumption behind it is broken.
The assumption is that engineers will register models. The reality is that the registry only sees what a human or CI pipeline checks into it. There is no force field at the inference layer that pulls a model into the registry when it starts serving traffic. A model can run for two years, score millions of predictions, and never appear in the catalog.
The deeper problem is ownership. Models are often created, deployed, and then forgotten, with no one tracking fit for purpose. No one claims a notebook model because no one is assigned to claim it. The data scientist who wrote it left the team. The engineer who deployed it does not know who owns it.
The registry has an owner field, but the field is empty because no one volunteered to fill it. This is the root cause of MLOps sprawl. The broader pattern is the same one that drives AI sprawl across the enterprise. Disconnected teams adopt their own stacks because no one owns the cross-cutting discipline. Different teams pick different vector databases, different prompt tools, different serving runtimes.
MLOps sprawl is what you find inside the platform team. AI sprawl is what the CTO sees from the boardroom.
The fix is not a better registry with stricter gates. Stricter gates only push more models into the shadow. The fix is a different kind of inventory, one that starts from inference traffic, not voluntary registration.
The Four-Layer Model Audit That Finds Every Shadow Model

The audit replaces the registry as the source of truth. It works from the outside in, never asking the registry for permission. Four layers, run in parallel, each surfacing models the others miss.
Layer 1 - Inference traffic analysis. Instrument the serving mesh, API gateway, or feature store to log every model call by artifact hash and endpoint. Run for 30 days. The result is a fingerprint list of every model that scored traffic in production, including the ones no one remembers deploying. This is the layer that catches the silent majority.
Layer 2 - Compute and spend analysis. Pull GPU-hour billing, notebook scheduler logs, and serverless inference invoices. Surface workloads that consumed compute but never appeared in any registry or deployment manifest. Compute consumption is mandatory. Registration is voluntary. Billing records often surface more deployed models than the registry tracks.
Layer 3 - Code archaeology. Grep the monorepo for serialized model artifacts. The patterns are stable across stacks:
1grep -r --include="*.py" -E "(joblib\.load|pickle\.load|onnxruntime\.load|torch\.load)" . | wc -l2grep -r --include="*.ipynb" -E "(\.pkl|\.joblib|\.onnx|\.pt|\.bin)" notebooks/ | head
Each match is a checkpoint. Each checkpoint is a model that left a notebook and entered a serving path. Map them back to notebook paths and you have a lineage that the registry never recorded.
Layer 4 - Stakeholder mapping. Interview product, finance, and ops. Attach a business owner to every unique fingerprint surfaced in layers 1-3. Without an owner, the artifact is a liability, not an asset.
Why does this method produce durable systems? Because it is artifact-centric, the same pattern the verified MLOps reference recommends for model artifact strategies. The inventory is built from what is actually running, not from what someone promised to register. Systems still serving five years after deployment share one trait: the inventory was real on day one, and it stayed real.
Once you know the real number, consolidation is shorter than the chaos suggests. Here is the 90-day path used across deployments in regulated industries.
The 90-Day Consolidation Playbook
The playbook is not a transformation program. It is a focused operation with three phases, each 30 days, each producing a tangible artifact leadership can review.
Days 1-30: Audit and baseline. Run the four-layer audit in parallel. Dedupe model fingerprints by hash, framework, and serving endpoint. Publish a baseline inventory to leadership. The number will be uncomfortable. That is the point. You cannot consolidate what you have not named.
Days 31-60: Triage and migrate. Sort every fingerprint into three buckets: retire, containerize, or register. Retire duplicates and zombie models that have not served traffic in 90 days. Convert the survivors to deterministic pipelines. Check them into the registry with full lineage, signed artifacts, and an owner. The gate looks like this in GitHub Actions:
1name: model-artifact-gate2on: [push]3jobs:4 registry-check:5 runs-on: ubuntu-latest6 steps: - uses: actions/checkout@v4 - name: Detect unsanctioned artifacts7 run: |8 if find . -name "*.pkl" -o -name "*.joblib" -o -name "*.onnx" | grep -q .; then9 if ! grep -q "registry_version:" .mlops/manifest.yaml; then10 echo "::error::Serialized model without registry metadata"11 exit 112 fi13 fi
The build fails if a serialized model ships without a registry entry. The gate is automatic. The lineage is enforced.
Days 61-90: Lock the path. Require every notebook export to flow through a CI job that registers the artifact, signs it, and attaches an owner. The notebook is no longer the production artifact. The artifact is the registered, signed, owned model that the notebook produced.
The 90-day window is realistic because each phase produces a concrete artifact. The audit produces a baseline inventory. Migration produces a curated registry. Governance hardening produces a gate that blocks untracked exports. Teams that skip the methodology and build the audit from first principles add time, because the failure modes of untracked model sprawl are discovered rather than inherited.
What Changes When You Actually Know What Is Running
Audit responses compress from weeks to hours. A regulator asks for an AI inventory, you query the artifact store, you have the answer. Compute spend drops once duplicate and zombie models are retired, since retired workloads stop consuming GPU hours and serverless inference cycles. Incident response becomes possible: when a model misbehaves, you can find every copy and roll it back.
The same discipline that closes the sprawl gap also makes the next model launch faster, not slower. Engineers stop reinventing inference scaffolding for every notebook export. The registry becomes a tool they want to use, because it removes the bureaucratic friction that pushed them into Jupyter in the first place.
Teams that fix the inventory once keep it fixed, because the playbook pays for itself every quarter. Our work with regulated enterprises on AI estate consolidation has shown the pattern holds from a single platform team to a multi-business-unit enterprise. The shift is not a one-time cleanup. It is the operating model that keeps the next 200 models from joining the shadows.
Frequently Asked Questions
Q: How do you find shadow ML models that never went through the registry?
A: You stop searching the registry. Instead, instrument inference traffic at the serving layer, audit GPU and serverless billing, grep the codebase for serialized model artifacts, and interview stakeholders. The four-layer audit catches models that were never registered because it never assumes the registry is the source of truth.
Q: What is the difference between MLOps model sprawl and AI sprawl?
A: MLOps model sprawl is a subset of AI sprawl. It focuses on the operational layer: models, pipelines, and serving infrastructure. Broader AI sprawl also includes disconnected tools, vector databases, datasets, and use cases across business units. MLOps sprawl is what you find inside the platform team. AI sprawl is what the CTO sees from the boardroom.
Q: Can Jupyter notebooks ever be production-ready?
A: Not directly. The verified research is explicit. Notebooks executed cell by cell are not deterministic. They often cannot be reproduced by their own authors. They lack the lineage a production model needs. The fix is converting the notebook into a scripted, versioned, CI-tested pipeline, then registering the resulting artifact, not the .ipynb file.
Q: How long does it take to clean up MLOps model sprawl?
A: The 90-day playbook has three phases, each 30 days, each producing a distinct artifact: an inventory, a curated registry, and an enforced gate. End-to-end, the timeline depends on estate size, stakeholder coordination, and the complexity of pipeline migration. Teams that build the audit methodology from first principles spend additional time discovering failure modes that a documented playbook already encodes.
Q: What tools detect shadow AI models in production?
A: Combine four signals. First, API gateway or service mesh logs for inference calls. Second, cloud billing exports for GPU and serverless usage. Third, code search for pickle, joblib, ONNX, and .ipynb checkpoints. Fourth, lightweight stakeholder surveys. No single vendor tool covers all four, which is why a layered audit outperforms any off-the-shelf scanner.
About the author
Mayank Singh is a software developer at Levitation Infotech, where he builds web and AI-powered applications across the company’s fintech, healthcare, and enterprise projects.
