A finished spot exists once. Its bytes, though, exist a dozen times: on the online NAS, in a project folder someone zipped for a client, in a cloud bucket a producer spun up for a review, on last year's LTO set, and in the deep archive nobody has touched since delivery. Every one of those copies is byte-identical. Every one of them pays rent every month. Most of the time nobody knows they are there.
This is the quiet tax on media storage. Not the working files you need, but the duplicate media you already paid to keep and then paid to keep again. It rarely shows up as a line item, because it looks exactly like real data. The NAS fills, IT buys another shelf, and the true cause, redundancy you cannot see, never gets named.
Post and studio workflows manufacture duplicates as a side effect of doing the work. A conform pulls a source local. A finishing artist snapshots a reel before a change. An I/O coordinator stages a deliverable, ships it, and leaves the staged copy behind. A vendor drop lands twice because the first Aspera session stalled. None of these are mistakes. They are the normal texture of a busy pipeline, and each one leaves a full-resolution copy sitting somewhere.
The trouble is that no single tool sees all the places at once. Your NAS has its own view. Your cloud console has another. Your tape catalog lives in a third system entirely. So the same 400 GB show file can sit on four tiers and read, to each system, as four legitimate and separate things. Nobody is wrong. Everybody is paying.
The reason duplicate media hides so well is that most tools look at the wrong thing. A filename lies constantly: the same edit gets renamed on the way out the door, and two genuinely different plates can share a name across shows. Comparing by name and size gets you a guess.
Gateway, the read-only console over your whole media estate, compares by checksum. It reads across every store, cloud, on-prem and LTO tape included, and returns files with their checksums and metadata. When two objects hash to the same value, they are the same bytes, full stop, regardless of what they are called or which tier they sit on. That is what makes media deduplication reporting trustworthy: it is an identity match, not a name match.
Gateway runs that comparison across 8 storage tiers and produces a checksum-level duplicate report alongside a storage-cost report. It is worth being precise about the word "dedup" here, because two different products do two different jobs. Gateway reports duplicates in the catalog so you can decide what to reclaim. When you are ingesting new material, Nexus is the one that verifies checksums on intake and can skip a source-path duplicate before it ever lands. Reporting on what exists is Gateway. Refusing to write another copy is Nexus. You want both, and they are not the same button.
A duplicate count is interesting. A dollar figure is what moves a budget conversation. Gateway's cost report is priced on your storage rates, per tier, so the output is not a generic industry estimate, it is what those redundant copies actually cost you this month. That matters because reclaim value depends entirely on where the copy lives. A duplicate sitting on hot online storage is expensive. The same duplicate on cold tape is close to free to keep and rarely worth the restore effort to remove. A credible report tells you the difference instead of flattening it.
Note: Gateway is read-only by design. A duplicate report never deletes anything. It hands your team the checksum evidence and the cost, and the decision to reclaim, and on which tier, stays with you.
When the NAS hits ninety percent, the reflex is to buy capacity. It is fast, it is a known quantity, and it makes the alert go away. It also grows the exact problem you never measured, because the new shelf fills with the same mix of real work and invisible duplicates as the old one. You are renting more space to store copies you already own.
This is not an argument against ever buying storage. Real libraries grow, and cheap cold tiers are genuinely the right home for a lot of finished work. It is an argument for knowing which is which before you sign the purchase order. If a third of what you are about to pay to store is a copy of something one tier away, that is a number you want in the room.
Read-only, across cloud, on-prem and LTO tape. It reads checksums and metadata in place. Nothing is copied out and nothing is moved.
Objects that hash to the same value are grouped as true duplicates, spanning all 8 storage tiers, whatever they are named on each system.
The cost report applies your per-tier rates, so an expensive online copy reads differently from a near-free tape copy. You see reclaimable cost, not a raw byte count.
Your team holds the evidence and the dollar figure. What comes off which tier is your call, made on numbers instead of a hunch.
Duplicate reporting is one report Gateway runs, and it sits next to the rest of what it does: deep search that returns files with checksums across the estate, catalog Q&A with permissions resolved per question, and the RBAC, MFA and SSO front door with a severity-tagged audit trail. The same read-only console that finds the copies is the one your IT and security team already trusts to look but never touch. If you also want the archive itself to stop being a black hole, the companion piece on why tape is a feature, not a graveyard covers the search-and-preview layer that makes cold storage browsable.
The fastest way to see your own duplicate media number is the four-week pilot: read-only, on one production or facility and archival data you choose, with the findings delivered in your own numbers. It is scoped to answer exactly the question this post raises. How much of what you store, you already store somewhere else.