Reference folders accumulate duplicates in a handful of predictable ways: a browser tab you'd already saved and didn't recognize the second time, an old low-resolution export sitting next to a full-resolution one added later, or an entire folder copied wholesale during a move between drives and never cleaned up afterward. None of these are Coii Ref problems to solve, exactly — they happen upstream of anything a reference manager does, in how files get onto your disk in the first place. What matters for this page is what to actually do about them once they're there, given that the tool you're searching with isn't going to point them out on its own.
There's no dedupe scan, no hash comparison, no "merge duplicates" menu item. That's not an oversight — it's the same design that keeps files where they already are instead of copying them into a managed library, and a tool that never copies your files has less reason to also police them for copies. What Coii Ref can do is help you find the candidates faster once you know what you're looking for, and make sure that whatever you decide to delete disappears from search cleanly afterward.
It's worth being upfront about this before describing a workaround, because the honest fix here is sometimes "use a different, purpose-built tool for this specific step, and use Coii Ref for the search and tagging that comes after." A dedicated duplicate finder that hashes file contents will always be more accurate at the narrow job of "are these two files byte-identical" than anything built primarily for search — that's not a gap worth apologizing for, it's a different tool for a different question.
1. Separate the two kinds of duplicate you're actually dealing with
Getting this distinction right up front saves the most time of anything on this page, because reaching for the wrong tool for either kind wastes an afternoon without actually solving the problem. A hash-based dedupe tool run against near-duplicates that differ by even one pixel of crop will report them as entirely different files and miss them completely; a by-eye comparison of a folder that's actually full of exact byte-for-byte copies is a slow way to do something a hash comparison would finish in seconds.
"Duplicate" covers two different problems and they need different tools.
Byte-identical copies — the same file saved twice, maybe with a (1)
appended to the filename — are a job for a dedicated duplicate-finder
utility that hashes file contents; nothing in a metadata-based reference
tool, ours included, can do that reliably. Near-duplicates — the same shot
exported at two resolutions, or downloaded from two different sources at
slightly different crops — are a judgment call no hash comparison would
catch anyway, and that's the half of the problem a query line actually
helps with.
1.5 Know when a "duplicate" is actually two different assets
Before deleting anything, check that the two files you're comparing are actually the same asset and not two related-but-distinct ones that happen to look similar — a before-and-after pair, two frames from the same shot a few seconds apart, or a full image alongside a tight crop someone made of it deliberately. These aren't duplicates in the sense this page means, even though a quick glance might flag them as candidates. When in doubt, a note explaining the relationship is safer than a delete you can't undo, and it costs you nothing to keep both if you're not certain.
2. Clear exact duplicates in Finder first, before you touch tags
Run a dedicated dedupe pass over the raw folder before it's absorbed into your workspace, or on an import folder before you fold it into an existing one. Deleting the exact copies at this stage means you're never rating or tagging the same image twice under two filenames, which is the more tedious version of this problem to clean up later.
~/Pictures/Reference/Incoming- Downloads batch
- Duplicates removed
- coii-thumbnailswritten by Coii Ref
- coii-proxieswritten by Coii Ref
- coii.dbwritten by Coii Ref
3. Let the watcher clean up after you, not before
Once you've deleted the exact copies, don't do anything else. The watcher notices the removal and the file drops out of the index on the next scan — its tags, rating and notes go with it. There's no separate step to purge a dangling entry, because there isn't one to begin with.
4. Hunt near-duplicates by dimensions, since a hash won't catch them anyway
Two exports of the same shot at different sizes are common enough in a reference folder — a low-res version saved from a search result years ago, a full-res one added later — and a width filter surfaces the pairs worth comparing by eye:
width:>3000 height:>3000 likely the full-resolution keeper
width:<800 likely an old low-res resave
tag:duplicate flagged, decision made, not yet deleted
5. Tag first, delete in a batch
Rather than deciding and deleting one at a time, tag the ones you're cutting
as tag:duplicate and keep working. When you're ready, pull up the tag and
clear the whole batch in Finder at once:
Deciding calmly and executing in a batch beats deleting file by file while you're mid-search for something else entirely, which is usually when duplicates actually get noticed. This is close to the same rating-and-batch discipline that works for culling a folder down to your best shots — deciding first, acting once, rather than making a hundred small decisions under time pressure.
6. Prefer a note over a delete when you're not sure
If two versions of a shot might genuinely be worth keeping — a wider crop
alongside a tighter one, say — a note explaining the difference is safer
than a guess you'll regret. has:notes tag:duplicate finds the ones you
flagged but talked yourself out of deleting outright, which is a real and
common outcome of this whole process.
6.5 Decide your threshold for "close enough" before you start
Not every near-duplicate is worth resolving. Two exports of the same shot that differ by a handful of pixels in crop, or a JPEG re-save that's visually identical to the original, are rarely worth the time to compare side by side and pick a winner — the cost of keeping both is a few kilobytes and one extra row in a search result; the cost of hunting down and resolving every one of them by hand is your afternoon. Reserve the tag-and-batch workflow for duplicates you can tell apart at a glance — genuinely different resolutions, or files you know for certain are identical — and let the merely-similar ones sit. A library doesn't need to be duplicate-free to be useful; it needs the duplicates it does have to not get in the way, which a rating and tag filter already handles without requiring them to be deleted at all.
7. Watch for duplicates arriving through the importer, not just Finder
If you're migrating from another app, the importer carries across ratings, tags, notes and canvases by matching relative paths against your existing folders — it doesn't create new copies of your images, since your originals were never moved or rewritten in the first place. Duplicates that show up right after a migration almost always predate it: they were already sitting in the source library as two files, and the importer faithfully brought across two sets of metadata along with them. Running the dedupe pass in step 2 on the old library's export folder, before you point Coii Ref at it, catches this earlier than discovering it after the fact.
8. Treat near-duplicates from multiple sources as a tagging problem, not just a deletion one
Sometimes the "duplicate" isn't really redundant — it's the same reference
saved from two different places because you found it twice, weeks apart,
and didn't recognize it the second time. Rather than always deleting one,
consider whether the two versions actually carry different value: one might
have better colour, the other a note you wrote the first time that's worth
keeping. tag:duplicate has:notes surfaces exactly this case — pairs
you've flagged as redundant but where at least one carries context worth
preserving before you decide which file to keep.
9. Accept that a perfectly clean library is not the goal
It's worth naming the actual objective here, because it's easy to lose track of it partway through a dedupe pass: the goal isn't zero duplicate bytes on disk, it's a library where a search doesn't get meaningfully worse because of them. Those are different standards, and the second one is far easier to hit. A few dozen leftover near-duplicates in an old folder you rarely search don't cost you anything real; the same duplicates concentrated in a folder you search every week are worth the ten minutes the tag-and-batch workflow takes. Let the actual cost — not a completeness instinct — decide where you spend the cleanup time.
What to do when the duplicates are thousands deep across years of folders
At real scale, chasing duplicates after the fact stops being worth the time
it costs — the leverage is at the front door, not the back. Run the dedupe
pass on anything new before it's absorbed into the workspace, keep a
Downloads staging folder separate from your organized one, and accept that
a small residue of near-duplicates in an old folder is not worth an
afternoon to hunt down by hand. The tag-and-batch workflow above is for the
occasional cleanup pass, not a permanent chore — most of a forty-thousand
image library's duplicates arrived in the first two years, before any of
this discipline existed, and that's fine. Search still works around them;
they're just extra rows in a result you're already filtering by rating and
tag.
The math is worth being honest about, too: a library with a few hundred
duplicate files scattered across forty thousand costs you almost nothing in
search quality — a query for tag:practical-light rating:>=4 returning one
extra near-identical row out of sixty is a rounding error, not a problem
worth an afternoon of hunting. Where duplicates actually hurt is when
they're concentrated — an entire old export folder duplicated wholesale
into a new one — and that specific case is worth finding and fixing exactly
once, rather than trying to keep the whole library duplicate-free as an
ongoing standard nobody could realistically maintain across years of
collecting.
If your existing tool copies files into its own library rather than indexing them in place, that's usually where duplicates start multiplying in the first place — does Eagle copy your files covers that specific mechanism, and the full comparison with Eagle covers what changes about duplicate management when a tool stops copying and starts indexing where files already sit. For getting the rest of a large, messy folder under control at the same time, backing up a reference library properly is worth reading alongside this one — a dedupe pass and a backup pass are naturally the same afternoon of work. And if the pile you're deduplicating is genuinely large — tens of thousands of files across several drives — indexing a forty-thousand-image folder covers the first pass itself, which is worth doing before rather than after a dedupe sweep.
Worth trying against a folder you already suspect is full of duplicates — the trial is thirty days with search, tags and notes all included, free to run the width-filter workflow above against your actual mess before deciding it's worth cleaning up by hand.
Point it at the messiest folder you have rather than a tidy one — that's the folder this workflow is actually meant for, and thirty days is enough time to run a real dedupe pass on it, tag what's left, and see whether the result is meaningfully more useful to search than it was before you started. If it isn't, you've lost nothing but the time the pass took; if it is, you'll have a clean sense of how much of your existing duplicate problem is actually worth solving versus how much was never costing you anything to begin with.