2026-10-02 · 17 min · computer-vision · object-detection · vision-transformers · tensorrt · benchmarks · inference-optimization · explainer
A DETR predicts a fixed set of boxes in one shot, matches them to the ground truth with a bipartite assignment, and skips non-maximum suppression entirely. That design has been elegant since 2020 and slow since 2020. RF-DETR (Roboflow — Robinson, Robicheaux, Popov, Ramanan, Peri; arXiv 2511.09554, ICLR 2026) is the version that is both elegant and fast: a DINOv2 vision-transformer backbone wired to a deformable-attention decoder, and — the part worth reading the paper for — an end-to-end weight-sharing neural architecture search that trains one network from which every published size is drawn. RF-DETR-2XL is, by the authors' claim, the first real-time detector to pass 60 AP on COCO.
Object detection is where I spend a lot of time, so the thing I want to pin down is not the headline — it is what the headline is made of. The named sizes, N through 2XL, are not six models someone trained six times. They are six points someone picked off a single continuous Pareto curve, because the search space is baked into one training run and the operating point is chosen afterward without any fine-tuning. That is the genuinely reusable idea here, and it is also why some of the numbers move depending on where you read them.
| Repo | roboflow/rf-detr · latest release v1.11.1 (2026-09-30) |
| Paper | arXiv 2511.09554 v2, "RF-DETR: Neural Architecture Search for Real-Time Detection Transformers" (ICLR 2026) |
| Project | rfdetr.roboflow.com |
| Tasks | detection, instance segmentation, keypoint detection (preview) — one API |
| Backbone | DINOv2 ViT (replaces LW-DETR's CAEv2) |
| Sizes | N, S, M, L, XL, 2XL detection; N-2XL Seg; keypoint preview |
| License | rfdetr package + N/S/M/L detectors + all Seg + keypoint = Apache-2.0; XL/2XL detection (rfdetr_plus) = PML 1.0 |
| Headline | 2XL hits 60.1 COCO AP50:95 at 17.2 ms (T4, TensorRT, FP16, batch 1) |
- license
- Apache-2.0
- branch
- develop
- tests
- 195 files
- source
- 7.8 MB
- commit date
- 2026-09-30
by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded
local clone, 2026-10-02 at 8913f7b — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount
shallow clone: counts describe the pinned tree, not the history
What a DETR is, and why "real-time DETR" was an oxymoron
A classical detector — the YOLO family, Faster R-CNN — predicts a dense grid of candidate boxes and then runs non-maximum suppression to collapse the pile of near-duplicates into one box per object. NMS is a hand-tuned, non-differentiable post-process with its own threshold, and its latency is data-dependent: a crowded frame is slower than an empty one. The Detection Transformer (DETR, 2020) threw that out. It carries a fixed number of learned object queries, lets them attend to image features through a transformer decoder, and emits exactly that many boxes. Training uses the Hungarian algorithm to find the one-to-one matching between predictions and ground-truth objects that minimizes total cost, so each query learns to claim at most one object. No grid, no duplicates, no NMS. Detection becomes direct set prediction.
The catch was speed. The original DETR needed hundreds of epochs to converge and ran nowhere near real time. The line of work that fixed this — Deformable DETR (sparse attention to a few sampled locations instead of dense global attention), then RT-DETR, LW-DETR, and D-FINE — chipped the decoder and the attention pattern down until a DETR could run in single-digit milliseconds on a T4. RF-DETR sits at the end of that line and acknowledges its parents directly: it is built on LW-DETR, DINOv2, and Deformable DETR. The decoder still uses deformable cross-attention; what changed is the backbone in front of it and the way the whole thing is sized.

The structure is worth reading left to right. Patches plus positional embeddings go into the ViT. Inside each backbone block, a couple of windowed encoder layers (self-attention restricted to a local tile of tokens, cheap) alternate with a non-windowed layer (global attention, expensive but mixes information across the whole image). The backbone's multi-scale features pass through a projector to the decoder, where learned queries run self-attention among themselves, deformable cross-attention into the image features, and a feed-forward network, repeated across decoder layers. Class and box heads turn each surviving query into a label and a rectangle. A second lightweight head of depthwise convolutions reuses the same features to paint instance masks. Losses are applied at every decoder layer, not just the last — which is what makes it safe to drop decoder layers at inference time, a knob we will get to.
The DINOv2 backbone: borrowing internet-scale priors
LW-DETR, the model RF-DETR modernizes, used a CAEv2 backbone. RF-DETR swaps it for DINOv2. This is not cosmetic. DINOv2 is a vision transformer self-supervised on a very large curated image corpus, and initializing the detector's backbone from those weights, rather than from an ImageNet-scale classification backbone, is what lets the model generalize to datasets that look nothing like COCO. The paper's motivating observation is that specialist detectors quietly overfit COCO through bespoke architectures, learning-rate schedules, and augmentation schedules, and then fall over on real-world data with a different number of objects per image, a different class count, or a different imaging modality. An internet-scale pretrained backbone is the antidote: it arrives already knowing what surfaces, textures, and object boundaries look like.
There is a cost, stated plainly in the paper. CAEv2's encoder has 10 layers at patch size 16; DINOv2's has 12. More layers, more latency. The authors' answer is that they "make up for this latency using NAS" — the search finds configurations that claw the time back elsewhere. The ablation is clean: in the backbone study (Table 6), DINOv2 outperforms CAEv2 by 2.4% AP under identical pretraining and hyperparameters. Interestingly, SAM2's Hiera-S backbone, despite fewer parameters than SigLIPv2, came out considerably slower — Hiera's speed claim does not survive contact with the Flash-Attention kernels TensorRT compiles for a plain ViT. If you want a real-time detector, a vanilla ViT you can fuse beats a cleverly-structured backbone you cannot. Foundation-model vision backbones are becoming the default front end for detection the way they already are for retrieval and depth — the same bet SigLIP 2 makes on the language-image side.
Two training details make the DINOv2 transplant actually work. The learning rate drops to 1e-4 (LW-DETR used 4e-4) with a per-layer multiplicative decay of 0.8, so the early backbone layers — the ones holding the most general features — barely move and DINOv2's pretrained knowledge is preserved rather than washed out. And the multi-scale projector uses layer norm instead of batch norm, which is what lets the whole thing train on consumer GPUs via gradient accumulation instead of needing a batch large enough for batch-norm statistics to be stable.
Weight-sharing NAS: one supernet, every size
Here is the center of the paper. Classical neural architecture search is expensive because it trains a candidate, measures it, and repeats — and hardware-aware variants repeat the whole search for each new target device. RF-DETR borrows the trick from Once-For-All (OFA): decouple training from search by training a single over-parameterized network — a supernet — whose sub-networks all share weights. At every training iteration you sample a random architecture configuration and take a gradient step on that configuration. Because the weights are shared, every sub-network is being trained a little, all the time. The authors liken it to ensemble learning with dropout, and note a free side effect: this "architecture augmentation" acts as a regularizer and improves generalization on its own, to the point that introducing weight-sharing NAS improves even the base configuration — despite that base using patch size 14, which is not in the search grid.

The search space is five knobs, shown above: input resolution, patch size, number of decoder layers, number of query tokens, and the number of windows per windowed-attention block. The crucial property is that the first — a trained supernet — is never evaluated until it has been fully trained on the target dataset. After that, every sub-network already performs well, so finding the accuracy-latency frontier is a grid search over configurations with no retraining. Resolution is interpolated by pre-allocating positional embeddings at the largest resolution over the smallest patch and interpolating down. Decoder layers and queries are simply dropped at inference — safe, because the loss was applied at every decoder layer during training. The paper evaluates 6,468 network configurations at search time: 11 resolutions × 7 patch sizes × 7 decoder-layer counts × 3 window settings × 4 query settings.
What does this cost, honestly? Training a NAS supernet runs roughly two to four times as long as training one non-NAS baseline, and the architecture search itself is estimated at about 10,000 GPU-hours (200 T4 GPUs for 48 hours). That is not free. But the comparison is not "2XL versus one YOLO" — it is "the entire N-through-2XL ladder from one run" versus "retrain a separate model for every size and every latency target you care about." For a team that actually ships several operating points — a nano model on a camera, a large one in the cloud — a single training run that emits the whole curve is the cheaper bill. And the paper reports that sub-networks never explicitly sampled during training still perform well, so the curve is continuous, not just defined at the six named dots.
Checking the 60 AP claim
The headline: RF-DETR-2XL is the first real-time detector to surpass 60 AP on COCO. The numbers behind it, read from the repo's README benchmark table — where every row is re-measured in Roboflow's own SAB harness with pycocotools over the full 5,000-image val2017 split, on a T4 with TensorRT, FP16, batch 1, so the comparison is apples-to-apples:
| Model | COCO AP50:95 | Latency (ms) | Params | Res | License |
|---|---|---|---|---|---|
| RF-DETR-N | 48.4 | 2.3 | 30.5M | 384 | Apache-2.0 |
| RF-DETR-S | 53.0 | 3.5 | 32.1M | 512 | Apache-2.0 |
| RF-DETR-M | 54.7 | 4.4 | 33.7M | 576 | Apache-2.0 |
| RF-DETR-L | 56.5 | 6.8 | 33.9M | 704 | Apache-2.0 |
| RF-DETR-XL | 58.6 | 11.5 | 126.4M | 700 | PML 1.0 |
| RF-DETR-2XL | 60.1 | 17.2 | 126.9M | 880 | PML 1.0 |
| D-FINE-X | 59.3 | 11.5 | 62.0M | 640 | Apache-2.0 |
| YOLO26-X | 56.9 | 9.6 | 56.9M | 640 | AGPL-3.0 |
The claim holds. At 60.1 AP, the 2XL clears 60; the strongest comparison points, D-FINE-X at 59.3 and YOLO26-X at 56.9, do not. The paper's own Table 2 reports the 2XL identically at 60.1 AP and 17.2 ms. "First real-time detector past 60 AP" is a defensible, checkable statement.
Two honest caveats sit next to it. First, the jump from XL (58.6) to 2XL (60.1) is 1.5 AP for 5.7 ms more latency and essentially the same parameter count (126.4M to 126.9M) — the 2XL mostly buys its accuracy with resolution, running at 880×880 against the XL's 700×700. Second, and this is the one that matters if you are choosing a model to ship: the two sizes that cross into Apache-territory accuracy are not Apache-licensed. N, S, M, and L — everything up to 56.5 AP — are Apache-2.0. XL and 2XL detection ship only in the rfdetr_plus extension (pip install rfdetr[plus]) under PML 1.0. The number that broke 60 is the one you cannot use under a permissive license. D-FINE-X, at 59.3 and Apache-2.0, is the model to beat if your lawyers are in the room. (The segmentation and keypoint heads invert this: all Seg sizes, including Seg-2XL at 49.9 mask AP, and the keypoint preview at 71.8 AP, are Apache-2.0.)

The Nano margin, and why the paper and the README disagree
The second claim to check: RF-DETR-Nano beats D-FINE-Nano by 5.3 AP at similar latency. In the paper's Table 2, RF-DETR-N is 48.0 AP at 2.3 ms; D-FINE-N is 42.7 AP at 2.1 ms. 48.0 minus 42.7 is 5.3, at latencies that are within 0.2 ms of each other. That checks out, and it is a large margin for a nano-class model — RF-DETR-N at 48.0 matches what YOLOv8 and YOLOv11 need their medium sizes to reach.
But notice the number already drifted. The README table above lists RF-DETR-N at 48.4, not the paper's 48.0. This is not an error in either; it is the two documents measuring different artifacts at different times. The README says so explicitly: its COCO numbers are re-measured in-house in the SAB harness, computed with pycocotools over all of val2017, and "may differ from vendor-reported figures." The paper reports the configuration as it stood at submission; the README reports the released v1.11.1 checkpoints re-run through the current harness. The drift is small and it lives at the low end — Nano moves 48.0 to 48.4, Small moves 52.9 to 53.0 — while M, L, XL, and 2XL match to the decimal. If you are citing a number, cite the one whose protocol you can see: the README states its harness, hardware, precision, and batch size on the page. Against the README's own 48.4, the margin over D-FINE-N widens to 5.7; against the paper's 48.0, it is the stated 5.3. Both are real; they are just not the same experiment, and the honest move is to carry the protocol with the number.
RF100-VL, and the "20×" that is really a runtime mismatch
The third claim lives on RF100-VL, Roboflow's 100-dataset benchmark built to measure exactly the out-of-distribution generalization COCO cannot. Here the comparison is not against another specialist but against a fine-tuned vision-language detector: RF-DETR-2XL versus GroundingDINO-tiny. The abstract states 2XL wins by 1.2 AP "while running as fast." The numbers: in the paper's RF100-VL evaluation, GroundingDINO-tiny scores 62.3 AP at 309.9 ms, and RF-DETR-2XL (fine-tuned) scores 63.5 AP at 15.6 ms. 63.5 minus 62.3 is the stated 1.2 AP. (The README re-measures the 2XL at 63.2 AP50:95 — the same low-end drift as on COCO.)
Now the speed. 309.9 divided by 15.6 is 19.9 — so "about 20× faster" is where that figure comes from, and it is in the data. The caveat is that it is not a clean like-for-like speedup, and the paper is careful not to headline it as one. GroundingDINO's 309.9 ms is a PyTorch measurement — the paper marks it with a star because that model does not support TensorRT execution — while RF-DETR's 15.6 ms is TensorRT. You are comparing an unoptimized runtime against an optimized one. The honest reading is the paper's own: RF-DETR matches a fine-tuned VLM's accuracy on real-world data "at a fraction of the runtime," and the fraction is large but part of it is that the VLM cannot be compiled the way the specialist can. The underlying point survives the caveat intact: a 127M-parameter specialist with an internet-scale backbone can match a 173M-parameter VLM on out-of-distribution detection while being deployable in real time, which a heavy text-encoder-bearing VLM is not. This is the same specialist-versus-foundation-model tension that runs through SAM 3.1's throughput story and the training-free SAM3-to-detector conversion in DART — the field keeps rediscovering that a focused model you can fuse beats a general one you cannot.
The measurement discipline is half the contribution
A detail I did not expect to respect as much as I do: the paper spends real effort on the fact that detector latency benchmarks are mutually incomparable, and ships a tool to fix it. Each new model re-benchmarks its predecessors on its own hardware, and the numbers wander — D-FINE's reported latency for LW-DETR is 25% faster than LW-DETR's own paper reported. The authors trace this to GPU power throttling during back-to-back forward passes, and find that simply pausing 200 ms between passes stabilizes the measurement. They also catch the common sin of reporting accuracy in FP32 and latency in FP16 — two different artifacts — and show that naive FP16 quantization can drop a model to near 0 AP (D-FINE fell to 0.5 AP until they fixed its ONNX export to opset 17). Their rule: report accuracy and latency from the same artifact, and they release the standalone harness (SAB) so anyone can reproduce the rows. For someone who ships models, this is as valuable as the architecture. A benchmark you cannot reproduce is a number you cannot trust, and most detection benchmarks are the former. The deployment story — TensorRT, FP16, CUDA graphs to pre-queue kernels — is the same production-glue layer that decides whether real-time perception actually runs on the robot, not just in the paper.
What I'd take from reading it
The weight-sharing NAS is the idea I would actually reuse. A single supernet, trained with architecture augmentation so every sub-network is a working model, from which you grid-search an accuracy-latency frontier without a single retraining run — that is the right shape for any team that ships more than one operating point of the same detector. Resolution, depth, queries, patch size, and windows are the knobs that happen to transfer, and the paper's ablations on which knobs matter for which dataset characteristics are the part I would read twice before adapting this to a new domain. The DINOv2 transplant is a clean demonstration that internet-scale backbones belong in real-time detectors, latency cost and all, as long as you keep the backbone fuseable.
The three claims all hold, with the right footnotes attached: 60.1 AP is real but PML-licensed; the 5.3-AP Nano margin is real at the paper's protocol and a hair larger at the README's; the 20× is a real ratio that compares a PyTorch runtime against a TensorRT one. None of those footnotes is a gotcha — they are the difference between a number and a number you can ship against, and the paper and repo are unusually candid about every one of them.