Skip to content

Repository files navigation

Local AI Image Detector

A Chrome extension that flags AI-generated images on the pages you visit. Every image is decoded and classified inside your browser. No cloud inference, no API, no localhost helper, nothing uploaded — the extension's only network request is for the image the page had already loaded.

MIT licensed, including the model weights.


Install

Requires Node 18+ and Chrome 116+.

git clone https://github.com/agentatwork/local-ai-image-detector
cd local-ai-image-detector
npm install        # onnxruntime-web
npm run build      # vendors the runtime, downloads + verifies the model weights

Then open chrome://extensions, turn on Developer mode, choose Load unpacked, and select the repository directory.

npm run build is the only step that touches the network. After it finishes the directory is self-contained: disconnect entirely, restart Chrome, and the extension still works. It downloads one file — the model weights — from the pinned Hugging Face revision named in tools/model.json, and checks its SHA-256 before writing it. A build either reproduces the exact bytes these measurements were taken against or fails.

Use

Browse. Images at least 128px on a side are scored as they scroll into view, and every one of them gets a badge in its top-left corner carrying what it scored — AI 91% in red at or above the 65% threshold, real 88% in grey below it. Click the toolbar icon to badge only the flagged ones instead, change the threshold, set the minimum size, or turn scoring off.

The threshold is 0.65, which is above one half, so there is a band where an image is not flagged and is still more likely generated than not. Badging that real 37% would be arithmetically true and would read as a contradiction, so the band reads unsure 63% AI instead. Whatever the badge says, the number is on the scale the word names, and the tooltip always gives the raw probability of AI.

Scoring every image out loud is the default on purpose. If only the flagged images are badged, a real photograph and an image the extension never looked at look exactly the same to you, and a silent detector is indistinguishable from a broken one — which is a failure mode this model demonstrably has, on small JPEGs, where it degrades by going quiet rather than by getting things wrong.

To ask about one specific image — including one below the size cutoff, or with automatic scoring switched off entirely — right-click it and choose Check this image for AI generation.


How it works

content.js     finds images in view        (IntersectionObserver, no page reflow)
     |  url
background.js  routes and caches by URL    (holds no model)
     |
offscreen.js   fetches + decodes           (one model instance for all tabs)
     |  ImageBitmap
preprocess.js  cap short side to 768px     (Lanczos, and only if it is larger)
     |  ImageBitmap
preprocess.js  two views, 384x384 each     (own bicubic; the canvas is not asked)
     |  tensor x2
detector.js    ViT -> mean -> calibrated probability

Inference is a ViT-Small/16 at 384px — Community Forensics (CVPR 2025), trained on 2.7M images from 4,803 different generators, which is why it holds up on generators it has never seen. It runs under WebGPU where the browser offers it and WebAssembly everywhere else, via onnxruntime-web.

The model lives in an offscreen document rather than the service worker (which Chrome kills mid-inference and which has no WebGPU) or the content script (which would mean one copy of the model per tab).

Three things that matter more than the model

Calibrate the threshold. Raw model output is not a probability. This one is nearly certain about real images and only mildly confident about generated ones, so reading it at a fixed 65% cutoff turns a good detector into a high-precision, low-recall one and throws away most of its accuracy. tools/calibrate.py fits a two-parameter Platt scaling that puts the measured decision boundary at 0.65. It is a monotone rescaling: it changes no ranking and no AUROC, only the number a reader sees.

Score every image twice, and not with two crops. The obvious intuition is that resizing destroys the evidence, since this classifier reads resampling artefacts, so you should crop at native resolution and never scale. Measured, that intuition is half wrong: the native crop is worse on its own than the model card's downscale, and nearest-neighbour upscaling of small images — which preserves the pixel grid perfectly — is worse still, because it does not preserve evidence so much as manufacture it, and it manufactures the same evidence for real photos as for generated ones (FFHQ-256 goes from 17% false positives to 92% — measured against the previous view pair and never re-run, since that view was rejected before squash shipped).

What works is two views that disagree in useful ways. The pair that ships is native (a 384px centre crop at the image's own resolution, which reads quantisation) and squash (the whole frame bicubic-resized to 384×384, aspect abandoned, which reads composition). An earlier build shipped native alongside the model card's downscale-then-crop, official; squash displaced it, and the reason is not the one I expected. I assumed squash would win only on heavily downscaled images, where cropping into an already-small frame sees almost nothing. It wins on every condition, undegraded included — 87.5% against official's 84.1% on pristine files. Both of those are single-view balanced accuracy with each view read at its own best threshold, measured on the 320-image sweep as it stood before the input cap shipped, and official is not a view this build loads any more: they are the evidence for a decision, not a description of what runs. The distinguishing property is not robustness, it is that a squash is the only view that never throws pixels away; both crops discard everything outside the crop window.

The price of that is aspect distortion, and it shows up where you would expect it: on the 49 images in the sweep with an aspect ratio of 1.40 or worse, squash's margin over official narrows from +3.3 to +1.2 points, the same rule refitted inside the stratum. Forty-nine images cannot establish a two-point difference-of-differences, so read that as consistent with the explanation rather than as evidence for it. Which views are averaged lives in model/config.json, and a calibration fitted for one set is meaningless applied to another, so the two travel together.

Do the resize yourself. ctx.drawImage into a smaller canvas resamples with whatever filter the compositor picks, which is neither documented nor stable across Chrome releases — and the choice of filter is part of what the model is reading. preprocess.js implements Pillow's bicubic explicitly: same support scaling, same normalisation, same rounding between passes. tools/compare.py checks that the shipped JavaScript and the Python that fitted the threshold agree on real images.


Measurements

The eval set is 1,020 images the model has never seen: 540 generated by 18 different generators, 480 real from 13 public datasets plus 90 random files pulled off Wikimedia Commons. Balanced accuracy throughout, read at the shipped 0.65 threshold.

Every figure below comes from tools/verify.mjs — the built extension, in a real headless Chrome, with the browser switched offline. All 1,020 images scored, none errored. Since the 768px input cap shipped, the run they are read from is capship/verify_cap.json; the uncapped verify_all.json is kept beside it as the comparison.

balanced accuracy recall (AI) specificity (real)
shipped: two views, calibrated 85.9% 85.9% 85.8%
leave-one-generator-out 86.9%
squash alone, best threshold 85.6%
native crop alone, best threshold 79.3%
two views, uncalibrated, read at 0.65 67.7%

AUROC is 0.9324, and calibration does not change it — it is a monotone rescaling. What it changes is where 0.65 lands: on the raw scale the shipped boundary sits at 0.003728, so reading the model's own output at a fixed 65% cutoff costs 18.2 points.

That 85.8% specificity is a property of this eval set, whose real images stop at a 1280px short side. On full-resolution camera originals the same build flags far more of them — measured, with a three-arm control, in Out of distribution: large camera originals below. Read that before quoting the number here, and Input normalization for how much of that gap a short-side cap closes and what it costs on the AI side.

Split by size, the two-view average wins in both strata, which is why it is not a composition artefact:

all min side ≥384px min side <384px
native alone, best threshold 79.3% 89.3% 66.8%
squash alone, best threshold 85.6% 88.9% 81.6%
mean of the two, best threshold 87.0% 92.0% 80.7%
mean of the two, shipped 0.65 85.9% 90.3% 80.4%

The last two rows are the price of shipping a fixed operating point rather than an oracle one: one point of clean accuracy, paid to have a threshold that was chosen without seeing this set. The shipped threshold is not fitted here at all — it is fitted by minimax over the eleven delivery conditions in tools/minimax.py, which is why it is not the best row. tools/views_clean.py regenerates this table from the same capship/verify_cap.json.

Nearest-neighbour upscaling of small images was also tried and is worse than either (78.3% on a balanced subsample, measured against the previous view pair — it was rejected before the squash view replaced the downscale, and was not re-run). It recovers exactly the generators the downscale misses — GenImage BigGAN 0% → 100%, FLUX.1-dev at 256px 0% → 75% — and simultaneously drives FFHQ-256 from 17% false positives to 92%. It does not preserve evidence; it manufactures it, equally for both classes.

Where it fails, honestly: leave-one-generator-out, the two worst held-out generators are GenImage BigGAN (50.9%, recall 6.7%) and a 256px FLUX.1-dev subset (54.3%, recall 13.3%) — a GAN and a small-resolution diffusion crop, both at chance. Nine of the eighteen generators score above 95%; the full fold list is printed by tools/report.py.

Leave-one-generator-out comes out at 86.9%, slightly above the 85.9% headline, and that is not a held-out gain — each fold refits its threshold on clean images, while the headline is read at the fixed threshold this extension actually ships, which was fitted by minimax over delivery conditions and never saw this set. The gap is the price of one fixed boundary.

Python/JavaScript parity. tools/verify.mjs scores images through the built extension in a real headless Chrome with the browser switched offline; tools/compare.py diffs that against the Python that fitted the calibration. On 128 images — every 8th of the eval set, so the stride walks both classes and all 31 sources rather than taking a prefix — the two implementations agree to a median of 3.8e-07 on the mean of views, worst case 4.1e-03, and zero decisions change at the shipped threshold. That comparison predates the input cap; the capped build's equivalent is H1 in Making it: the cap, in the browser, run over all 1,020 images rather than 128.

The two views disagree by very different amounts, and the reason is worth stating because it is a check on the implementation rather than a curiosity: native agrees to a median of 4.9e-09 because it does no resampling at all — it is a pure crop, so both sides are reading the same pixels and only the normalisation arithmetic can differ. squash resamples the entire frame and lands at 7.0e-07, worst case 8.2e-03. The residual is Pillow's fixed-point coefficient rounding against JavaScript's float, and it appears exactly where resampling happens and nowhere else, which is what it should do if the bicubic was transcribed correctly and not what it would do if the two were running different algorithms.

The headline numbers above come from a separate full-set run through the same harness, so they are the extension's own scores rather than the Python's. Parity is checked on the stride because this box has one core and a full-set Python re-scoring costs hours; the stride answers the question parity actually asks, which is whether the two implementations agree image by image.

Speed: about 3.3 s per image, two views, on the single WASM thread of a one-core cloud VM. WebGPU and real hardware are considerably faster; the extension only scores images that scroll into view, and caches by URL.


Privacy

  • No image data leaves the device. There is no code path in this repository that sends image bytes, pixel data, features, hashes or scores anywhere.
  • The only requests made are GETs for image URLs the page has already fetched, without credentials.
  • Nothing is written to disk except your settings.
  • <all_urls> host permission exists so the extension can read cross-origin images. A page-context canvas cannot read them back, which is why the fetch happens in the extension.

Reproducing the numbers

Every script that produced a number in this README is in tools/. The eval images themselves are not redistributed here — the fetchers rebuild the set from public sources.

pip install onnxruntime pillow numpy
python3 tools/fetch.py real bitmind/MS-COCO 30   # build an eval set from public datasets
python3 tools/fetch_web.py 90                    # plus real images off the open web
python3 tools/variants.py 12                     # which preprocessing keeps the evidence
python3 tools/dump.py logits.json native squash   # score the set the way the extension does
python3 tools/calibrate.py logits.json           # fit the Platt slope on the clean set
python3 tools/perview.py --views official,native,squash  # every view x every delivery pipeline
python3 tools/minimax.py                         # refit the intercept on the WORST pipeline
python3 tools/shiptable.py                       # the eleven-condition table as it ships
npm run build && node tools/verify.mjs data/ai verify.json --offline
python3 tools/compare.py logits.json verify.json # Python vs shipped JavaScript
node capship/h0_check.mjs                        # the input cap's resampler against PIL
python3 device/analyse.py                        # is the residual false positive a phone photo?
node tools/demo.mjs 4 demo.png                   # badges on a real page, screenshotted

tools/verify.mjs scores through the extension's code but never renders a page. tools/demo.mjs is the other half: it serves a grid of eval images over local HTTP, loads the unpacked extension, and waits for content.js to badge them of its own accord. Nothing in the page tells it which images are which.

tools/minimax.py writes the chosen views and calibration to minimax.json; those three values are then set in tools/model.json, which is the source npm run build writes model/config.json from. Editing model/config.json directly works until the next build and then silently reverts — a mistake I made once with views, caught only because detector.js refuses to start without them, and then made again with input_cap. model/ is not in this repository at all, so a value that only exists there never reaches a clone either. detector.js now refuses to start without input_cap too, for the same reason: the specificity this README quotes was measured with the cap in place.

tools/verify.mjs is the one that counts. It loads the built extension into a real Chrome, switches the browser offline, and scores images through the extension's own code path — so the reported figures come from the shipped JavaScript, not from the Python that chose the model.

Where it broke, and what fixed it

The headline above is measured on pristine dataset files. Images on the web are not pristine, so the same build was re-scored over 320 stratified images (18 generators x 10, 14 real sources x 10) through eleven delivery pipelines. Nothing changes but the pipeline: same weights, same two views, same frozen calibration, same 0.65.

The 768px input cap ships, so this table has been re-run through it. What that replaces is a hedge published in its place — the cap only acts on images with a short side over 768px, which the 320 stratified images here largely are not, but "largely" is not a measurement. Measuring it took two minutes against the image headers: 82 of the 320 are over the cap. Re-running the table took an hour.

pipeline resized by the cap old views v1.0 v1.1 recall (AI) specificity (real)
rescale 90%, no re-encode 80/320 86.6% 88.6% 88.3% 91.7% 85.0%
nothing 82/320 85.7% 86.2% 86.1% 87.2% 85.0%
JPEG q75 82/320 83.1% 85.8% 85.7% 85.0% 86.4%
CMS resize ≤640px + q80 0/320 81.5% 85.6% 85.6% 86.1% 85.0%
JPEG q90 82/320 84.7% 84.2% 84.1% 86.1% 82.1%
≤768px + JPEG q60 0/320 79.4% 83.3% 83.3% 76.7% 90.0%
JPEG q60 82/320 79.1% 81.8% 82.2% 74.4% 90.0%
CMS resize ≤1600px + q85 81/320 82.1% 81.7% 81.9% 83.9% 80.0%
CMS resize ≤1024px + q85 71/320 82.1% 81.7% 81.7% 83.3% 80.0%
WebP q80 82/320 79.7% 81.2% 80.8% 87.2% 74.3%
≤512px + JPEG q40 0/320 72.3% 79.0% 79.0% 67.2% 90.7%

Three numeric columns, three different operating points. old views is the previous build — official+native, threshold fitted on undegraded images. v1.0 is the shipped view pair native+squash at the minimax threshold, no cap. v1.1 is that same configuration with N(768) in front of the views, which is what the extension now does to every image. The second column counts the images the cap actually resized after the delivery pipeline had run: 642 image-condition pairs across eight of the eleven, and none at all in the three whose own resize already puts every image under the cap.

The cap changes almost nothing here — and that is the answer, not a reason for not having measured it. No condition moves by more than 0.4 points in either direction, the mean over the eleven is 83.5% before and after, eleven of eleven still clear the bounty's 75.0% bar, and the worst pipeline is still ≤512px + JPEG q40 at 79.0%, identical to the last digit because that pipeline's own resize makes the cap a no-op on all 320 images.

The direction of the small moves is worth naming, because it is the opposite of the result that motivated shipping the cap. On the eight touched conditions the cap mostly nudges recall up and specificity down — on undegraded images recall goes 86.7% → 87.2% and specificity 85.7% → 85.0%, one image each way out of 180 and 140. That is a fact about this corpus rather than a reversal: 67 of the 180 generated images have a short side over the cap against 15 of the 140 real ones, so here the cap is mostly acting on the generated half. The study that motivated it was measured on large camera originals, of which this stratified set contains fifteen. Same function, different images.

Read left to right rather than top to bottom, the table still says what it said before the cap existed: the pipeline that used to fail the bar at 72.3% scores 79.0%, and its recall — the number that was actually broken — goes from 51.1% to 67.2%. That failure mode was the point. Degraded hard enough, the old build stopped calling things generated rather than calling them wrongly, which from outside looks exactly like a detector working correctly on a set of real photographs.

This was bought, not found, and the price is still on the clean row. Specificity on undegraded images was 91.4% on the previous build and is 85.0% on this one — roughly one extra false positive per sixteen real photographs, in the condition an extension meets most often while you browse. The threshold was chosen to maximise the worst pipeline rather than the average one, and that is a deliberate answer to how the bounty is written: 75.0% is a floor, and the images it will be judged on have been through a delivery path nobody described to me. If you wanted the best average instead, you would pick a different intercept and get a quieter extension that fails harder on small recompressed images. Both are defensible; only one of them is the criterion here, and I have stated which one this is tuned for rather than presenting it as a free win.

Two numbers used to keep the win in proportion, and one of them was quoted at the wrong operating point. This section said AUROC at ≤512px + q40 is 0.841 against 0.927 clean. Those two values are right — for the previous view pair, inside a paragraph about the build that ships. AUROC is threshold-free, which is why it is quoted at all, but it is not view-free, and re-running the table is what surfaced that:

views AUROC, ≤512px + q40 AUROC, undegraded gap
previous build (official+native) 0.8410 0.9274 0.0864
v1.0 (native+squash, no cap) 0.8605 0.9393 0.0788
v1.1 (the same, capped) 0.8605 0.9402 0.0798

The argument survives the correction and gets slightly weaker: at the operating point that actually ships the pair is 0.8605 against 0.9402, a gap of 0.0798 rather than 0.0864. A real part of the loss at that pipeline is ranking loss that no threshold recovers. The cap cannot touch that row at all and moves the undegraded one by 0.001. And the honest out-of-sample figure is not 79.0% but 76.6% — see the next section, which is about how that was validated and why the earlier version of this section refused to ship a change of exactly this shape.

Two things the average hides. Both are quoted below at the shipped threshold with the cap in place, and both are identical to the uncapped run — for a reason worth stating rather than leaving as a coincidence: ADM's ten images are 256×256 and BigGAN's are 128×128, so none of those twenty is over the cap, and N(768) hands every one of them straight through. Neither subset's recall moves in any of the eleven conditions under the cap.

  • GenImage's ADM subset falls from 90% to 50% between clean and JPEG q75 — four images out of ten, from one generator out of eighteen, and those four are the entire corpus-wide recall loss at that step rather than merely most of it: every other generator holds. It survives WebP q80 (90%) and a 90% rescale (100%) untouched, so this is JPEG quantisation specifically, not resampling. On the previous build the same subset fell to 10%, so this is one of the places the new view pair actually did work.
  • GenImage's BigGAN subset scores 0% in nine of the eleven conditions. A 2018 GAN against a diffusion-era model: a blind spot, not a degradation. The two exceptions are odd enough to name — WebP q80 reaches 10% and a 90% rescale reaches 50%, i.e. this build detects a GAN better after the image has been resampled than before. On ten images that is one and five images respectively and I would not build anything on it, but it is the opposite of the direction everything else in this table moves, so it is recorded rather than smoothed away. Both averages above include a class this build largely does not detect.

Data, per-image probabilities and the script that recomputes all of it: agentatwork/c143-survey · the capped re-run, its halts and its gate live in robustcap/ · discussion in Twenty-two detectors.

Three attempts to repair the eleventh; the third one shipped

Because the failure is a threshold sitting in the wrong place — the per-condition oracle reaches 77.3% on the same scores that operate at 72.3% — a calibration that knew how damaged an image was should be able to recover most of the gap. A browser cannot be told the condition, so it needs to estimate the damage from the decoded bitmap: no re-fetch of the original bytes, no DQT marker, no network request added to a privacy extension.

Blockiness works as that estimator. JPEG quantises each 8×8 block independently, so it leaves discontinuities on the block grid that are absent between columns inside a block. The ratio of mean across-boundary to mean interior luma gradient is ~1.0 for an image that has never been through a block transform, and rises as quality falls. Measured over the same 320 images (blockiness.py, ~20 lines of numpy, no model):

pipeline blockiness
rescale 90% 1.000
nothing 1.255
JPEG q75 1.258
resize ≤1024px 1.224
≤512px + JPEG q40 1.600

It rises from clean to q40 on 99.1% of images individually, AUC 0.889 as a clean-vs-degraded test, and a 90% rescale reads exactly 1.000 because resampling destroys the grid. So the feature is real, cheap and honest.

Both calibrations built on it are worse than the constant that ships.

  • Refitting the Platt intercept globally on the three mildest pipelines, generators held out: worst condition 72.3% → 69.6%.
  • Letting the intercept move with the feature, b = b₀ + b₁·(blockiness − 1), generators and conditions held out: the fit chooses t(d) = −2.66 + 2.50·d, the wrong sign, and ≤512px + q40 goes 72.3% → 65.8% (recall 32.2%). Leave-one-generator-out, worst condition averaged over folds: 64.5%, minimum 45.0%.

The diagnosis is worth more than the attempt. The three fit pipelines span blockiness 1.224–1.258 — a range of 0.034 — while the pipeline being repaired sits at 1.600. The slope was never identified in-sample; it was extrapolated 2.5× outside the range it was fitted on. Nothing in the fit set could see it, and the fitted value's own sign is therefore noise.

One confound was suspected and ruled out rather than assumed away: if real photographs arrive as JPEG and generated images as PNG, blockiness would be a class label in disguise. It is not — AUC of blockiness against the AI label is 0.496 / 0.509 / 0.474 / 0.479 across four pipelines (0.5 is chance), and the file extensions are balanced (170 of 180 AI and 140 of 140 real are JPEG).

A third repair was tried, and it produced the most useful negative result of the three.

The two views shipping at the time were both crops, so each showed the model a fraction of the frame at close to native scale — good for reading quantisation, which is exactly what heavy JPEG destroys. So a fourth view was added: squash, the whole frame bicubic-resized to 384×384, keeping no pixel grid at all and keeping composition instead. It costs no extra download, because it is the same model asked a different question.

Read on ranking alone it is the best single view by a clear margin, and it is the most compression-robust:

view clean ≤512px + q40
official (shipped then) 84.1% 74.0%
native (shipped, then and now) 81.7% 76.7%
up 78.3% 69.8%
squash (shipped now) 87.5% 79.2%

It is still not shipped, and the reason is a trap worth more than the view.

Views have to be combined, and there are two obvious ways: average the probabilities, or average the logits. detector.js averages probabilities — const mean = (p1 + p2) / 2. Averaging logits instead is the kind of choice that looks like a matter of taste. It is not. Scoring every subset of the four views both ways, with one threshold fitted on the clean condition only:

combination worst condition, logit mean worst condition, probability mean
official+native+up+squash 77.4% 69.8%
official+squash 77.1% 72.3%
native+squash 74.8% 75.4%
official+native (shipped then) 71.9% 72.7%

Same scores, same images, same threshold procedure — a 7.6-point swing on the four-view combination purely from which space the mean is taken in, and it reverses the ranking: the best combination in logit space is the worst in probability space. The check that proves it is arithmetic rather than a bug is that all four single-view numbers are identical in both tables (73.5%, 70.8%, 73.5%, 68.5%) — with one view there is no average to take.

In the space the extension actually uses, the best combination reached 75.4% against the shipped path's 72.3%. At the time I wrote: +3.1 points, inside the ±2.8-point standard error, so it is not a result. That sentence was wrong, and the third repair is now what ships. What follows is why, kept in this order because the reasoning matters more than the outcome.

The error bar was the wrong one. ±2.8 is the standard error of one balanced accuracy estimated once — the right bar for an absolute claim like "this clears 75%". But comparing two view pairs on the same images is a paired design: both see the same photographs and make correlated errors on them, so what governs the comparison is the standard error of the difference, which is smaller. Measured with a stratified paired bootstrap (20,000 resamples, AI and real resampled separately to hold the class balance): +3.0 on the worst pipeline, 95% CI +0.8 .. +5.4. The interval excludes zero. Using the unpaired figure as a floor for a difference is not conservative in a harmless direction — it throws away real improvements, and it threw away this one. Across all eleven pipelines the candidate wins on ten and eight of the eleven intervals exclude zero.

The fit objective was also wrong, and this mattered more than the error bar. The old threshold was fitted on undegraded images, which optimises the condition the extension is least likely to be in. Refitting it to maximise the worst of the eleven pipelines — the minimax choice, since the bounty's 75.0% is a floor — moves the worst pipeline to 79.0% rather than 75.4%.

Which raises the obvious objection: a threshold fitted on eleven conditions and then reported on those same eleven conditions is a memorisation score. It is, so that is not the number to trust. The honest one is leave-one-condition-out: fit the threshold on ten pipelines, score the eleventh, rotate.

worst condition mean clearing 75.0%
previous build 71.5% (WebP q80) 79.3% 9/11
v1.0, no cap 76.7% (≤512px + JPEG q40) 83.1% 11/11
v1.1, capped 76.6% (≤512px + JPEG q40) 83.1% 11/11

Held out that way the shipping build's worst condition is 76.6% and all eleven still clear 75.0%. This one is re-fitted per fold, so it is not readable off the table above and had to be re-run separately under the cap; the cap moves it by about a tenth of a point, in the same direction and for the same reason as the table rows. 76.6% against a 75.0% bar is a thin margin and I am not going to dress it up as more; it is, however, the number that answers the question, and it is the one this build is chosen on.

And the trap above still stands. Nothing here rehabilitates averaging in logit space — detector.js averages probabilities, and every number in this section is measured in that space, which is the only space the claim is about. The 7.6-point swing was real and the lesson survives the reversal of the conclusion.

Two things I want stated plainly, because each is a way this could be misread. I changed the test after seeing the data. That is the move that manufactures results, and naming it does not neutralise it; what makes this a fix rather than a rationalisation is that the paired test and the minimax objective are both more appropriate a priori, for reasons that have nothing to do with which answer they give. And "the improvement is real" is not "the result clears the bar." Those are two different questions taking two different error bars: the gain is established by the paired interval, while 76.6% against 75.0% is an absolute claim carrying the full unpaired ±2.8. The gain is established. The clearance is not — it is the best estimate I have, on a benchmark I have never seen.

The full argument, in a form someone else can run against their own detector rather than mine, is in PROTOCOL.md.

Out of distribution: large camera originals

When this was measured, the extension reported 86.46% specificity — a 13.54% false-positive rate — on the 480 real images in its eval set. It has since shipped a 768px input cap, which moves that figure slightly; the current one is in the headline table at the top of this README, and the study below is unchanged by it. On a different corpus, 598 camera originals pulled off Wikimedia Commons, the same build at the same threshold flags 36.62% of them. Both numbers are read at the identical shipped operating point, so the operating point is not the difference. The difference is size, and the eval set does not contain the sizes where the detector fails.

The eval set's largest real image has a 1280px short side. The Commons corpus's median is 2495px. Where the two overlap they agree closely; the eval set simply stops before the range in which the false-positive rate doubles.

short side of the original eval set (real) arm A arm B arm C
under 384px 25/213 = 11.7% none none none
384–768px 17/167 = 10.2% 5/49 = 10.2% 6/49 = 12.2% 6/49 = 12.2%
768–1500px 23/100 = 23.0% 21/84 = 25.0% 7/84 = 8.3% 13/84 = 15.5%
1500–3000px none 81/244 = 33.2% 31/244 = 12.7% 63/244 = 25.8%
3000px and over none 112/221 = 50.7% 36/221 = 16.3% 80/221 = 36.2%

The three arms

To find out whether size is causal or just a marker for "camera originals are different", the 598 originals were scored three ways. All three carry the same q97 re-encode, so the encode is not what separates them. This was pre-registered before any image was scored, bands and all.

what the detector saw flagged as AI Wilson interval
arm A the originals, exactly as the agreement study scored them 219/598 = 36.62% 32.86% – 40.56%
arm B resampled (Lanczos) to a 512px short side, re-encoded q97 80/598 = 13.38% 10.88% – 16.34%
arm C a 512x512 centre crop at native resolution, re-encoded q97 162/598 = 27.09% 23.68% – 30.79%

Arm B and arm C have the same short side and nearly the same pixel count. Only arm B has been resampled. So the pair separates the two explanations of a size effect: resampling smoothing the evidence away, versus the detector seeing too little of a large frame.

Note which way the confound points. Resampling removes sensor noise, and on this model's own logic missing sensor noise should read as more synthetic, not less. Arm B was predicted to fall anyway, against the direction the confound pushes — and it falls hard, to 13.38%, which is the eval set's own rate to within a point. Arm C falls too, to 27.09%, but only about a third as far. Of the 219 images arm A flags, arm B leaves 75 still flagged and arm C leaves 144. Both manipulations help and resampling helps far more.

The mechanism I predicted, and got wrong

The pre-registration said the size effect would live in one of the two views:

If size is causal, native carries it and squash does not.

native is a 384x384 centre crop at native resolution, so on a large photograph it covers a vanishing fraction of the frame. squash resamples the whole frame to 384x384, so its input does not depend on the original's size at all. That prediction fails:

view what it sees Spearman ρ with log(short side)
native a 384x384 centre crop at native resolution 0.300
squash the whole frame resampled to 384x384 0.309

Both views drift toward "AI" on larger originals by about the same amount, and squash drifts marginally harder — 0.309 against 0.300 — despite having no way to know how large the original was. Whatever the detector is reacting to survives being resampled to 384x384, so it is not the fraction of the frame the crop covers.

Taking the 219 originals arm A flags and watching both views move under each arm says where each manipulation bites. This decomposition was not pre-registered, and it is medians rather than means because the mean of these scores is carried by a handful of near-1.0 outliers — arm C's mean squash score is higher than arm A's while arm C flags 57 fewer images.

median native median squash still flagged, of 219
arm A 0.01929 0.00834 219
arm B 0.00016 0.00136 75
arm C 0.01776 0.00210 144

Arm C leaves the native view almost exactly where arm A had it, which is expected: it is the same 384x384 window on the same unresampled pixels. Everything arm C changes, it changes through squash — and that alone drops 75 of the 219 below the threshold, leaving 144. Arm B then collapses native by two orders of magnitude as well, and that is where the rest of the drop comes from. So the size effect is not one view's fault. The whole-frame downscale that squash performs is manufacturing part of it, and the detector's reading of unresampled large-sensor pixels is manufacturing the rest.

Every pre-registered band

pre-registered band measured
P1 spearman rho(native, log short side) on arm A 0.2 to 0.7 0.3004 in band
P2 spearman rho(squash, log short side) on arm A -0.2 to 0.25 0.3089 outside
P3 arm B (resampled to 512) false-positive rate 0.04 to 0.18 0.1338 in band
P4 arm C (512 crop at native scale) false-positive rate 0.26 to 0.45 0.2709 in band
P5 arm C minus arm B, in points 10.0 to 100.0 13.7124 in band

The bound written down before the run was:

If arm B's false-positive rate is above 30%, short side is not the driver and the size table above is a coincidence of these two corpora; the write-up says so and the mechanism paragraph is wrong.

It is not above 30%. Short side is the driver, the size table is not a coincidence of two corpora, and the part of the mechanism paragraph that fails is the which view claim, not the size claim. Four of the five bands hold; the one that misses is the view ordering, and it misses in the direction that makes the extension look worse rather than better.

What this means if you use this extension

The two sections below supersede this one, and a reader who stops here will act on the wrong number. Everything above was measured on a build that handed the model whatever resolution the image arrived at. The extension now caps its own input at 768px before either view runs, and on these same 598 originals that takes arm A's 36.62% down to 23.08%. What survives the cap is the shape of the problem rather than its size: a full-resolution file straight off a camera is still this detector's worst case, and it still gets worse as the file gets bigger.

What the section said when it was written, and what the next one was answering:

On a phone screenshot, a web-sized JPEG, or anything already resampled for the internet, the published specificity is the number that applies. On a full-resolution file straight off a camera, it is not: expect roughly a third of them to be called AI, rising with the size of the file. Downscaling such an image to a 512px short side before checking it moves the false-positive rate back to the published range — that is arm B, and it is a workaround rather than a fix.

Two things in that paragraph did not survive being measured. The workaround is no longer the reader's job, because the build performs it; and back to the published range was too generous, because arm B carries a q97 JPEG round-trip that the resize alone does not — the next section separates the two and finds the resampling worth about nine points less than this one credited it with.

The honest summary is that the 86.46% specificity measured here describes the eval set, and the eval set under-represents large camera originals badly enough that the number does not transfer to them. The fix belongs in the training and calibration data, not in the README, and the cap is not that fix: it is worth 13.5 points on camera originals and it is paid for in eval-set specificity, 86.46% down to 85.83%. A trade, not a correction.

Scoring code, pre-registration, full result artifact and the prose gate that checked this section: oob/. Ten of the 598 originals are already smaller than 512px on the short side, so for those arm B upscales and arm C has nothing to crop; the arms table is reported over all 598 either way, and the artifact carries the same rates over only the 588 the arms genuinely shrink.

Input normalization: what capping the short side actually costs

The section above ends on a workaround:

Downscaling such an image to a 512px short side before checking it moves the false-positive rate back to the published range — that is arm B, and it is a workaround rather than a fix.

That was only ever measured on real photographs. A preprocessing step that halves the false alarms by halving the detections is not an improvement, so this section runs the same normalization over the labelled eval set too and reports both sides at the identical shipped operating point: same weights, same two views, same Platt fit, same threshold. Only the pixels handed to the model change.

N(k) — if the short side is larger than k, Lanczos-downscale so it is exactly k; otherwise pass the image through untouched. It never upscales and it never re-encodes, because a browser extension resizes on a canvas, in memory. Two caps were pre-registered with bands, halts and a bound before any image was scored: N512 as the primary and N768 as a gentler secondary.

what the detector saw images it resized
Baseline the shipped pipeline, unchanged none
N512 short side capped at 512px, Lanczos, never upscaled 445 of 1020 eval, 588 of 598 Commons
N768 short side capped at 768px, Lanczos, never upscaled 284 of 1020 eval, 543 of 598 Commons

The trade

recall (540 AI) specificity (480 real) balanced accuracy Commons FPR (598)
Baseline 85.93% 86.46% 86.19% 219/598 = 36.62%
N512 83.70% 87.71% 85.71% 135/598 = 22.58%
N768 85.93% 85.83% 85.88% 138/598 = 23.08%

N512 does what it was built to do: it takes the Commons false-positive rate down to 22.58%. That is 14.05 points below the Baseline's 36.62%, and of the 219 originals the shipped build flags it clears 99 while newly flagging 15. It is not free. On the labelled eval set N512 gives up 2.22 points of recall to buy 1.25 points of specificity — 14 AI images lost against 2 gained, which is a real difference rather than noise.

N768 is the surprise. The gentler cap lands at 23.08% on the same originals, and the difference between the two caps is 21 images one way against 24 the other. That is nothing: p = 0.77. On the eval set N768 costs no measurable recall at all — 2 AI images lost, 2 gained, p = 1.0 — for 0.63 points of specificity that is also within noise. The secondary arm gets essentially the whole out-of-distribution benefit at none of the price, and it does it by resizing 284 of the 1020 eval images instead of 445.

384–768px 768–1500px 1500–3000px 3000px and over
Baseline 5/49 = 10.2% 21/84 = 25.0% 81/244 = 33.2% 112/221 = 50.7%
N512 5/49 = 10.2% 14/84 = 16.7% 48/244 = 19.7% 68/221 = 30.8%
N768 5/49 = 10.2% 14/84 = 16.7% 46/244 = 18.9% 73/221 = 33.0%

The cap bites hardest where the baseline was worst and leaves the smallest bucket's rate unchanged, which is what a normalization should look like rather than a global shift.

The number that was at risk

The pre-registration named it: the 262 AI images with a short side above 512px are the ones the shipped build handles best, and they are the ones N(k) touches.

recall on the 262 AI images over 512px Wilson interval
Baseline 252/262 = 96.18% 93.12% – 97.91%
N512 240/262 = 91.60% 87.61% – 94.39%
N768 252/262 = 96.18% 93.12% – 97.91%

N512 drops them from 96.18% to 91.60% — 4.58 points, twelve images. N768 does not move them at all. The loss is not spread evenly:

generator n median short side Baseline flags N512 flags lost
Deepfake-leonardo-stablecog 22 1024px 22 22 0
GenImage_MidJourney 30 1024px 30 24 6
JourneyDB 30 1024px 30 30 0
bm-aura-imagegen 30 1024px 22 22 0
bm-diffusion 30 1024px 30 29 1
bm-imagine 30 720px 30 30 0
bm-mobius 30 1024px 30 30 0
klingai-images 30 768px 30 30 0
nano-banana 30 1024px 28 23 5

Eleven of the twelve are MidJourney and nano-banana. Size does not explain that: seven of the nine sets share a 1024px median short side and are therefore downscaled by the same factor, and of those seven, three lose images and four lose none. What the cap removes is specific to a generator, not proportional to how hard the image was resampled.

Is this a fix, or a threshold change in disguise?

The pre-registration raised that question and then did not specify the comparison that answers it, so this one is post-hoc. If moving the shipped threshold on the unmodified pipeline buys the same thing for the same price, the normalization is doing no work. The threshold is moved on the Baseline until it matches N512 on one axis, and the other two are then comparable.

threshold recall specificity Commons FPR
N512, shipped threshold 0.6500 83.70% 87.71% 22.58%
Baseline, threshold moved to match N512's specificity 0.6787 84.81% 87.71% 34.45%
Baseline, threshold moved to match N512's recall 0.7188 83.70% 89.58% 31.44%

On the labelled eval set the threshold wins outright: at matched specificity the Baseline keeps more recall than N512 does, and at matched recall it keeps more specificity. If the eval set were the whole world, N512 would be strictly worse than turning one knob.

The Commons column is why it is not. Matching N512's operating point by threshold alone leaves the false-positive rate on camera originals at 34.45% and 31.44%. N512 reaches 22.58% at the same eval-set trade. The normalization is not a disguised threshold change: it moves a population the threshold barely reaches, which is exactly the population the previous section showed the eval set does not contain.

The band that missed, and what it says about the re-encode

Five of the six pre-registered bands hold. P1 does not: N512's Commons false-positive rate was predicted at 8% to 20% and measured at 22.58%.

pre-registered band measured
P1 N512 Commons false-positive rate 0.08 to 0.2 0.2258 outside
P2 N512 eval recall 0.78 to 0.87 0.8370 in band
P3 N512 eval specificity 0.86 to 0.93 0.8771 in band
P4 N512 eval balanced accuracy 0.83 to 0.9 0.8571 in band
P6 N512 recall on the 262 AI images above 512px 0.8 to 0.95 0.9160 in band
P5 N768 Commons FPR strictly between N512's and baseline ordering 0.2308 in band

That band was drawn around the previous study's arm B, which put the same geometry at 13.38%. The two differ in exactly one step — arm B carried a q97 JPEG round-trip and N(k) does not — on the same 598 files under the same rule, and the baseline reproduces those studies to the image (219 flagged, both times).

what the detector was given Commons FPR (598)
original bytes, no normalization (oob arm A, this study's Baseline) 36.62%
512px short side, no re-encode (this study's N512) 22.58%
512px centre crop then q97 (oob arm C) 27.09%
512px short side then q97 (oob arm B) 13.38%

So roughly 9.20 points of what looked like a resampling effect was the re-encode, and my band inherited the error by treating the round-trip as an incidental detail of the earlier harness. Recompression is a real part of why a web-sized JPEG reads as human to this model, and a canvas resize on its own does not reproduce it.

The bound

Written before the run:

If N512's balanced accuracy on the labelled eval set is below 0.84, the normalization is not shippable as a drop-in, and the write-up says that in place of recommending it.

N512 clears it at 85.71%, so it ships on the pre-registered terms. But the evidence points at the other arm. N768 gets the out-of-distribution false-positive rate to 23.08% with no measurable cost on either eval-set axis, and the only thing recommending the harder cap is 0.50 points of Commons false-positive rate that a paired test cannot distinguish from zero. Capping the short side at 768px is the change worth making.

Neither cap gets camera originals to the 13.54% false-positive rate the eval set reports, and neither should be read as one. The published specificity still describes images the eval set actually contains. What this measures is that a one-line preprocessing step closes a little under three-fifths of that gap, and what the harder cap costs to close another half-point of it.

Scoring code, pre-registration, full result artifact and the prose gate that checked this section: capnorm/. Both caps are provable no-ops below their own cap — every eval and Commons image already at or under k scores bit-identically to the Baseline — and the Baseline re-scored in Python reproduces the shipped Chrome build's decision on all 1020 eval images, with a worst probability difference of 0.0062. A threshold re-fit on the normalized scores is in the artifact and is in-sample, chosen on the same images it is scored on; it is not a shipped number.

Making it: the cap, in the browser

The section above ends by recommending a change to this extension. This section is about actually making it, and about the part of that which was not a one-liner.

The recommendation was measured in Python, with PIL. The thing you install is JavaScript in a Chrome service worker. Those are two different programs, and the only honest way to publish a number from one against the other is to check that they agree — so the halts below were written before any of the code was, in capship/PREREG.md, and each of them can stop the change from shipping.

halt what it asserts measured
H0 the JavaScript Lanczos reproduces PIL's, on raw pixels 0/255 worst channel over 55,848,960 channels in 24 images, reductions up to 5.42×
H1 Chrome reproduces the Python arm's decisions 0 decision differences, worst
H2 images at or under the cap are untouched, bit-for-bit 736 untouched, 736 scoring identically to the shipped build; 284 resized
H3 every eval image scored 1020 scored, 0 errors

The resampler is the whole risk

N(k) is three lines of arithmetic: if the short side is longer than k, scale so it is exactly k; otherwise return the image untouched. Nothing about that is hard. What is hard is that the study measured PIL's Image.LANCZOS, and this extension's two views run a hand-written, Pillow-compatible bicubic. Reaching for the filter already in the file, or for the browser's own drawImage scaling, would have produced a build that downsizes images and reports the study's numbers — a different experiment wearing this one's clothes.

So H0 compares the two resamplers directly, on raw pixels, with the model out of the loop entirely. h0_dump.py hands PIL an image and writes both the decoded input and PIL's output; h0_check.mjs imports the shipping src/preprocess.js and runs it over the same bytes. A disagreement can then only be the filter. The sample is drawn from both corpora on purpose: the eval set's originals barely exceed the cap, while the Commons camera originals go far past it, and Lanczos averages more input pixels per output pixel as the factor grows — a resampler can be right at one and wrong at the other.

Pillow's 8-bit resample is also not float arithmetic. Coefficients are normalized and then quantized to 22 bits, accumulation runs in int32 with a half added up front and an arithmetic shift down at the end, C's cast truncates toward zero rather than rounding, and the horizontal pass rounds back to 8 bits before the vertical pass reads it. Reproducing that exactly, rather than approximately, is what the shipping version does, and it gets there: zero difference on every one of 55,848,960 channels across 24 images, reductions from 1.07× to 5.42×, portrait and landscape.

The first implementation was the obvious one — accumulate in floats, Math.round at the end — and it is kept in h0_control.mjs and re-run, because a parity test that has never rejected anything is not evidence about the thing it passed. Re-running it is what makes the next sentence possible, and the next sentence is not the one I expected to write.

H0 would have passed it. The float version lands 1 level out of 255 away from PIL, and 1 is exactly what H0 pre-registered as acceptable. Not once, either: all 24 images differ, and 0 channels differ by more than 1 — a uniform one-level error, right where the halt's boundary was drawn. So the check that reads as having caught something would in fact have shipped it, and what actually rejected it was a stricter standard I applied afterwards by hand, without writing it down. The bit-exactness above is a stronger result than H0 asked for; H0 is a weaker halt than it looks. If you are reusing this scaffolding, that is the transferable part — the control is what tells you where your check's real boundary is, and it is worth running even when, as here, the implementation under test passes.

What the browser then does

With the resampler settled, H1 is the question the whole task exists for: does Chrome running the capped extension produce what Python running the study produced? It does, and more exactly than the halt required. Zero decision differences on 1020 images — and because every decision matches, the arms are not merely close, they are the same numbers.

what ran recall (540 AI) specificity (480 real) balanced accuracy
shipped the extension as published, in Chrome 85.93% 86.46% 86.19%
capped the same extension with input_cap: 768, in Chrome 85.93% 85.83% 85.88%
Python N768 the study's N768 arm, in Python with PIL 85.93% 85.83% 85.88%

The capped row and the Python row are identical because they are the same 464 recalls and the same 68 false positives, image for image. That is the sentence the previous section was not yet entitled to write.

The aggregate worst-case probability difference, 6.170e-03, is worth reading carefully, because it is not a measurement of the new code. It falls on an image the cap passes through untouched, where the extension does exactly what it did before and the disagreement is the pre-existing one between the JavaScript and PIL bicubic in the views — the same computation the tolerance itself was derived from, which is why the two match to every digit. Split the other way, the 284 images the cap actually resamples agree to 1.121e-03, five times better than the 736 it leaves alone. The new path is the quieter half of the build.

The cost of the cap on this corpus is a real one and it is small. Recall does not move at all; specificity goes from 86.46% to 85.83%, three real photos stop being flagged and six start.

images lost the AI flag gained it McNemar p
AI images 540 2 2 1.0000
real images 480 3 6 0.5078

A nine-image reshuffle out of 480 is not a result — McNemar puts it at p = 0.5078 — and it is not meant to be. The eval set is where the cap costs something; it is not where the cap was adopted. That was measured on 598 Wikimedia Commons camera originals, which this Chrome run does not contain, and where the false-positive rate falls from 36.62% to 23.08%.

median per image
shipped build, all 1020 3116.3 ms
capped build, all 1020 3125.1 ms
capped build, the 284 it resizes 3250.9 ms

Resampling before inference is not free, but next to a WASM forward pass on one core it is close: 3116.3 ms per image becomes 3125.1 ms, and even counting only the 284 images that are actually resized, 3250.9 ms.

The bound

Written before the run:

If the capped build's Chrome-measured balanced accuracy on the 1020 eval images falls more than 1.0 point below the shipped build's Chrome-measured balanced accuracy, the cap does not ship.

The shipped build scores 86.19% balanced accuracy and the capped build 85.88%, a fall of 0.3125 points. The bound allowed 1.0, and the Python study had predicted 0.31. The cap ships.

pre-registered band measured
P1 worst channel difference between the JS and PIL Lanczos, in levels 0 to 1 0 in band
P2 capped Chrome balanced accuracy minus the Python N768 arm's, points -0.5 to 0.5 0 in band
P3a capped Chrome recall minus the Python N768 arm's, points -0.5 to 0.5 0 in band
P3b capped Chrome specificity minus the Python N768 arm's, points -0.5 to 0.5 0 in band
P4 images left untouched by the cap exactly 736 736 in band
P5 median ms/image, as a fraction of the shipped build's at most 1.25 1.0028 in band

What this changes if you use it

If you already have this extension, update it. Nothing you have to do changes and nothing about how it decides changes — same weights, same two views, same Platt fit, same 0.65 threshold. What changes is that a large photograph is now bounded to a 768px short side before either view sees it, instead of being squashed from whatever your camera produced.

On the eval set that costs 0.63 points of specificity and no recall. On large camera originals — the population that eval set does not contain, and the reason for the change — it removes about a third of the false accusations. If you have ever pointed this thing at a full-size photo you took yourself and had it come back "likely AI", that is the case this is for.

Two things I would rather say here than let you find out: the cap lives in tools/model.json, not in the model/config.json the extension loads, because the build overwrites that file every time — an earlier version of this change was edited in the wrong one and would have silently reverted. And src/detector.js now refuses to load a config with no input_cap at all, rather than treating a missing one as "no cap". The specificity published above was measured with the cap in place, and a build without it is a different detector wearing these numbers.

Verification code, pre-registration, full result artifact and the prose gate that checked this section: capship/.

Limits

Worth saying plainly:

  • A confident score is not proof. Heavily edited photographs, AI-upscaled real photos and screenshots of generated images all sit in genuine grey area.
  • Small images carry less evidence. Below about 200px the score should be read as a hint.
  • A full-resolution file straight off a camera is the worst case, and it is not close. The 768px input cap takes the false-positive rate on 598 Wikimedia originals from 36.62% to 23.08%, against 14.17% on the eval set's own real images — better, not fixed. The last three sections are the measurement.
  • What is left over after the cap is not phone photography, which is what I expected and pre-registered. Splitting those same originals by the capture device in their EXIF, mobile is flagged at 19.58% and a dedicated camera at 24.37%, and mobile stays lower in three of the four size bands. Phones are the larger files here, not the smaller ones, so they carry that penalty and are still flagged less: device/.
  • Every detector degrades on generators released after its training data. This one degrades more slowly than most, which is the whole reason it was chosen, but it still degrades.

Licence

MIT — see LICENSE. The model weights are MIT from buildborderless/CommunityForensics-DeepfakeDet-ViT. onnxruntime-web is MIT.

About

Chrome extension that flags AI-generated images. All inference in-browser: no cloud, no API, no localhost backend.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages