~/satyajit

Restore and relight LoRAs: teaching a generator to start from your picture

mdjsonmcp

2026-10-06 · 20 min · video-generation · image-generation · diffusion · lora · fine-tuning · open-weights · explainer

Two LoRAs came out in the last week of September, and on the surface they have nothing in common. Lightricks' LTX-2.5 Restore IC-LoRA takes archive video, "low-resolution web rips, low-bitrate broadcast transfers, tape, sepia or black-and-white film scans", and renders the same shot clean and in colour. RunningHub's Qwen-Image-2.1 lighting LoRA takes a product pasted onto a background and, in the repost's words, removes reflections and glare, enhances specular highlights and adds natural shadows.

Underneath they are the same move. A model trained to generate from noise and text is shown a picture, clean, inside its own token sequence, and a small set of extra weights teaches it to treat that picture as the thing to redraw. This piece explains that move from the matrix up, then checks each release against what can be read without a GPU: the cards, the safetensors headers over HTTP range requests, the LTX-2 trainer and pipeline code, DiffSynth-Studio, and ComfyUI's loader.

One constraint shaped the work. The Lightricks repository is gated behind a click-through, and this box has no Hugging Face token, so its weights, its README and its example videos all return 401. The card text is public on the model page, the file sizes are public in the API, and the comparison clip is in the repost. Everything about the restore weights below is therefore reasoned from a byte count, and labelled so.

Lightricks/LTX-2.5-22b-IC-LoRA-Restore@ed04134 · snapshot 2026-10-06
repo size
1.74 GB
task
video-to-video
library
ltx
license
other
safetensors
1 shard
largest file
1.71 GB
files
9
downloads
1.4K
likes
51
gated
auto
languages
en
ic-loravideo-to-videovideo-restorationcolorizationarchive-footageltx-2.3ltx-2.5ltx-video

Gated; weights not read. The file is 1,711,643,914 bytes (measured from the Hub API), which matches rank 128 on all six attention modules of all 48 blocks, 855,638,016 parameters, to within 9,127 bytes of header (reasoned). Trained on LTX-2.3, tested on LTX-2.5 unchanged (reported). The validation sources also appear in training under other degradations (reported). LTX-2 Community License.

repo last modified 2026-09-29

repo size
174.8 MB
task
text-to-image
safetensors
1 shard
largest file
167.8 MB
files
5
downloads
0
likes
64
comfyuilora

Measured from the header: 448 BF16 tensors, rank 32 on seven linear layers in each of 32 blocks, 83,886,080 parameters, 1.18% of the 7,115,124,736-parameter denoiser. Metadata: trained on ModelScope with DiffSynth-Studio settings, learning rate 3e-05, 2,000 steps, trigger word pengyu. The card describes no data, metrics or licence of its own; the base model's is the non-commercial Qwen Research License.

repo last modified 2026-09-30

A LoRA is a low-rank diff

Fine-tuning a layer means changing its weight matrix WW, of size dout×dind_{out} \times d_{in}. LoRA (Hu et al., 2021) freezes WW and learns the change as a product of two thin matrices:

W′=W+λ⋅αr BA,A∈Rr×din,  B∈Rdout×rW' = W + \lambda \cdot \frac{\alpha}{r} \, B A, \qquad A \in \mathbb{R}^{r \times d_{in}},\; B \in \mathbb{R}^{d_{out} \times r}

rr is the rank, α\alpha a fixed scale chosen at training time, and λ\lambda the strength slider in ComfyUI. BB starts at zero, so the adapter starts as a no-op. The cost is r(din+dout)r(d_{in} + d_{out}) parameters per layer instead of dindoutd_{in} d_{out}. For a 4,096 by 4,096 projection at rank 128 that is 1,048,576 numbers against 16,777,216, 6.25% of the layer (reasoned). At rank 32 it is 262,144, or 1.56%.

The bet is that the change a task needs lies in a few directions per layer. Restoration and relighting are good candidates: the base model already knows what a steam locomotive or a glass bottle looks like under daylight. The adapter has to teach it something narrower: where to look for the layout, and what to do differently once it has.

Both files set α=r\alpha = r, so α/r=1\alpha / r = 1 and the strength slider is the only scale. The restore card states rank 128, alpha 128. The relight file has no alpha tensor, and DiffSynth-Studio's add_lora_to_model defaults lora_alpha to the rank; ComfyUI applies a scale of 1 to a file with no alpha (all measured from the sources).

What is in the two files

The relight header is 58,240 bytes of JSON, read with two range requests. It lists 448 BF16 tensors: an A and a B for to_q, to_k, to_v, to_out.0, img_mlp.gate_layer, img_mlp.proj and img_mlp.out in each of 32 blocks. Each attention projection is 4,096 square; the two MLP inputs go 4,096 to 12,288 and the output comes back. That is 2,621,440 parameters a block and 83,886,080 in all, and 8 + 58,240 + 2 x 83,886,080 is the file's size to the byte, 167,830,408 (measured).

The restore header cannot be read. What can be read is the base model. The LTX-2.3 transformer header lists 21,005,004,544 parameters and six attention modules in each of 48 blocks: video self-attention, video-to-text cross-attention, audio self-attention, audio-to-text cross-attention, and the two cross-modal modules, audio-to-video and video-to-audio (measured). The LTX-2 paper's own diagram shows them:

Two diagrams. On the left, one LTX-2 block as two parallel stacks: the video stack runs self attention, T2V cross attention, A2V cross attention and an FFN on the video hidden state; the audio stack runs self attention, T2A cross attention, V2A cross attention and an FFN on the audio hidden state; text hidden state feeds both cross attentions. On the right, a detailed view of the bidirectional audio-video cross attention, with Q, K and V projections, temporal 1D RoPE, scale and shift from the timesteps, and a gate.
The dual-stream LTX-2 block: video and audio each have self-attention, text cross-attention and a cross-modal attention, six attention modules in all, every one with to_q, to_k, to_v and to_out (LTX-2 paper, Figure 2).

The card says the restore LoRA "targets to_q/to_k/to_v/to_out". In the LTX-2 trainer that is a suffix match, and the trainer's configuration guide warns that the short patterns "will match all attention modules including attn1.to_k, audio_attn1.to_k, audio_to_video_attn.to_k, and video_to_audio_attn.to_k" (measured from the docs). At rank 128 over all six modules of all 48 blocks, that is 855,638,016 parameters, 1,711,276,032 bytes of BF16. Rebuild the header those tensors would need, with the licence text the ungated Union-Control LoRA carries as metadata, and the predicted file comes to 1,711,634,787 bytes. The real file is 1,711,643,914: 9,127 bytes apart, less than one part in 180,000 (reasoned). Video attention alone would be 805,306,368 bytes, half the file, so the short patterns are what was used.

That has a consequence the card does not mention. The restore run trained on video; the card lists "Audio: Not trained for audio generation". With no audio in the batch, the LTX-2 transformer skips the audio and cross-modal branches entirely (run_ax, run_a2v and run_v2a are false), so their LoRA B matrices get no gradient and keep their zero initialisation. If that is what happened, 452,984,832 of the 855,638,016 parameters, 53%, are zeros that load, cost memory, and change nothing (reasoned from the trainer code; not checked against the weights).

The calculator below rebuilds every one of these files from the measured layer shapes. Pick a base, tick the layer groups, slide the rank. It reproduces the two ungated LTX-2.3 IC-LoRA headers I read (Lightricks' Union-Control at rank 64, 327,155,712 parameters, and a community colourizer at rank 32, 163,577,856, both on video attention and the video feed-forward) and both Qwen-Image-2.1 LoRAs on this site exactly.

base model

48 blocks · 21,005,004,544 base parameters

tick the layer groups the adapter targets; each bar is r x (d_in + d_out), summed over the group's layers and all 48 blocks

LoRA parameters855,638,016BF16 file, tensors only1,711,276,032 bytes (1632.0 MiB)share of the base4.07%
match this is the Restore IC-LoRA (gated): 855,638,016 parameters (reasoned)
Layer shapes measured from safetensors headers. The calculator reproduces the measured headers exactly; the Restore row is reasoned, because its file is gated and only its byte size is public. On Qwen-Image-2.1 the MLP input is either two 12,288-wide layers or one fused 24,576-wide layer; they are the same weights, so ticking one clears the other.

On Qwen the same layer budget can be spent two ways. Qwen-Image-2.1's MLP input is one fused gate_up weight in ComfyUI and ai-toolkit, but two separate 12,288-wide layers in diffusers and DiffSynth-Studio. ausboss's outpaint LoRA targets the fused one: one rank-32 adapter, 79,691,776 parameters. The relight LoRA targets the halves: two independent rank-32 adapters, 4,194,304 parameters more (measured). ComfyUI's loader has a branch for exactly this, mapping gate_layer and proj onto the two halves of the fused weight, and accepts PEFT's lora_B.default.weight naming, so the file loads as shipped (measured from comfy/lora.py and weight_adapter/lora.py at commit d49e888).

In context: the reference rides in the sequence

A LoRA changes weights. It does not, by itself, give a text-to-video model a way to see an input video. That is what the "IC" adds. The name comes from Alibaba's In-Context LoRA, which tiled related images into one canvas; Lightricks' version is more direct, and the LTX-2 trainer spells it out in training_strategies/video_to_video.py (measured from the code):

  1. Encode the target clip and the reference clip with the same video VAE. The VAE compresses 32x in height and width and 8x in time, and the patch size is 1, so a 960 x 544 x 97 window is 30 x 17 x 13 = 6,630 tokens.
  2. Noise only the target. The reference stays clean, and its per-token timestep is 0.
  3. Concatenate the two into one sequence and give both the same position grid. At downscale factor 1, reference token (t, h, w) and target token (t, h, w) carry the same rotary coordinates.
  4. Run the transformer over the lot, with ordinary bidirectional self-attention.
  5. Compute the flow-matching loss on the target tokens only.

At inference VideoConditionByReferenceLatent does the same thing: it appends the reference tokens with a denoise mask of 1 − strength, so a guide strength of 1 keeps them clean at every step. The order differs from training, reference first there and appended here, which attention does not care about: it sees positions, not order (reasoned).

That is the whole mechanism. There is no encoder for the reference and no new input channel. The target attends to the reference in every layer of the same network, and because the two share coordinates, "the matching pixel in the archive clip" is one rotary rotation away from every target token. What the LoRA learns is how to use that path: which parts of the reference to copy (layout, motion, structure) and which to override (grain, cast, blur, the absence of colour).

canvas

the trained tile

frames per window
one sequence, 13,260 tokens
target
archive clip
  • target: 6,630 tokens · noisy, sigma from the sampler, loss here
  • archive clip: 6,630 tokens · clean, timestep 0, no loss
keysqueries1111
latent grid30 x 17 x 13tokens in linear layers2.00x the target aloneattention pairs4.00x the target alonepositionsreference token (t, h, w) sits on target token (t, h, w)
Token counts reasoned from the LTX-2 code: a 32x, 32x, 8x VAE and patch size 1. The restore card runs 960 x 544 tiles of 97 frames with the clip at downscale factor 1, so the reference is as long as the target. The mask rule is read from ltx-core's build_attention_mask; the reference image's token count is an assumption (one latent frame of the canvas).

The cost is plain in the token count. At downscale factor 1 the reference is as long as the target, so every linear layer processes twice the tokens and self-attention four times the pairs: 13,260 tokens for one 960 x 544 tile of 97 frames (reasoned). Lightricks offers a cheaper variant elsewhere: the Union-Control LoRA's metadata says reference_downscale_factor 2, a quarter of the reference tokens, with positions scaled back up to the target's grid. The restore LoRA's card says factor 1, which makes sense for a task whose whole point is detail at the target's scale.

There is also a dial between full and no attention. Wrap a reference in ConditioningItemAttentionStrengthWrapper with strength s below 1, and build_attention_mask writes s into the target-to-reference blocks, 1 inside each reference, and 0 between two different references. The float mask becomes an additive bias of log⁡s\log s on the attention logits, so s = 0.5 halves the unnormalised weight of every target-reference pair (measured from mask_utils.py and transformer_args.py). At s = 1 no mask is built at all.

Qwen-Image-2.1 edits the same way with a different mask. Its reference image is a prefix in one stream of 32 blocks, modulated from t = 0, and the mask is block-causal: the target reads the reference, the reference never reads the target, as the release piece covers.

An attention mask drawn as a grid of query rows against key columns, divided into four segments: system prefix, optional input image, edit instruction, target image. The text segments are causal staircases; the image segments are solid blocks, bidirectional inside themselves. The target image rows see every column.
Qwen-Image-2.1's mixed-granularity mask. For the relight LoRA, the pasted composite is the input image and the relit picture is the target (Qwen-Image-2.1 release, architecture figure).

So the relight LoRA is in-context too, in everything but name. The difference that matters is in the arrow: LTX's reference and target see each other, Qwen's reference is read-only, which is what lets Qwen cache it.

How the restore LoRA was trained

The card is unusually complete about this, and it is the part to read before using it.

Synthetic pairs. "A clean 1080p clip was turned monochrome or tinted (neutral broadcast, cool tape, warm sepia), blurred as a period lens would, downscaled to 240p to 360p, encoded at a low bitrate one to three times over, sometimes denoised, scaled back up, and given film damage." The degraded clip is the reference and the clean clip the target. There are 151 training pairs and 8 held out, 97 frames at 24 fps, one shared caption (reported).

The recipe. The LTX-2 trainer's flexible strategy: the degraded video always present as the reference, the first frame given clean with probability 0.25, and the first 17 frames given as a prefix with probability 0.5. Rank 128, alpha 128, the Prodigy optimiser at learning rate 1.0, a cosine schedule, batch 1, 4,000 steps, buckets of 960 x 544 at 97 and 49 frames (reported). The prefix training is what lets a long clip be restored as chained 97-frame windows, each starting from the previous window's last 17 restored frames.

The settings it wants. Eight steps on the distilled model, guidance 1.0, LoRA strength 1.0, and the clip upscaled with lanczos to the working canvas before it becomes the guide, so "the model never sees the small scan". The canvas is cut into overlapping 960 x 544 tiles, the trained size, fused at every step, even for a full-HD output. The distilled sigmas are 1.0, 0.99375, 0.9875, 0.98125, 0.975, 0.909375, 0.725, 0.421875 and 0.0, so five of the eight steps run above 0.9 (reasoned): most of the schedule is spent where layout and colour are decided. Strength "behaves as a switch: at 0.7 the restoration turns off and the clip stays monochrome; 1.4 over-tints" (reported).

The card's gallery is three public-domain clips: a 1954 family meal, the Lumière train and the Hindenburg. Two of them, as they appeared in the repost:

A side-by-side crop. Left, labelled SOURCE 640x480, upscaled: a soft, grainy black-and-white frame of a boy in a dark sweater leaning over a table with a fork, a girl's hair in the foreground. Right, labelled RESULT 2880x2176, 1:1: the same pose in colour and sharp, a dark teal sweater, an orange wall, skin tones, individual strands of hair, a fork with food on it.
The 1954 family meal: the scan lanczos-upscaled on the left, the restored and refined result on the right, a 1:1 crop of the 2880 x 2176 output. The wall's orange, the sweater's teal and every strand of hair are the model's (Lightricks Restore IC-LoRA model card, gallery clip).
A side-by-side crop. Left, labelled SOURCE: the grey ribbed nose of the Hindenburg against a grey sky, with blurred trees and a mast below. Right, labelled RESULT: the same airship in blue-silver with crisp panel ribs, a blue sky, green trees and a yellow lattice mooring mast.
The Hindenburg at its mooring mast in 1937: the same crop before and after. The yellow mast is a colour decision, made from content and the prompt (Lightricks Restore IC-LoRA model card, gallery clip).

The numbers. On 8 validation pairs at 1920 x 1056, 97 frames, stage 1 only: PSNR +2.0 dB over the degraded input, a Lab colour error from 24 to 16, and high-frequency energy from 61% to 85% of the target, against +0.8 dB and 75% for the previous version (reported). Read the fine print next to them: those 8 pairs' "sources also appear in training under other degradations". The model has seen those clean frames. On clips whose sources it never saw, the card says, it "does not beat the aligned input on pixel PSNR; judge those by eye". That is the honest version of the claim, and it is Lightricks' own.

A reference image helps where it exists. Real crops of a sign, a logo or a face, placed over their own region on a grey canvas and attached at frame index −1, raised the sign region by 4.6 dB and the faces by 4.0 dB against the truth on one street clip (reported). It is the one mechanism here that adds information instead of inferring it.

What "restore" means here

Count what the model is given. The source is 640 x 480, one channel. The output is 2880 x 2176, three channels. That is 20.4 times the pixels, and two of the three colour channels have no source at all (reasoned). Everything between the scan's samples is drawn by a video generator conditioned on them.

The card says so in its own words. On the train clip, "the station side of the frame holds almost no information in the scan, so the model rebuilds it: the buildings, the hillside and the porter's cart on the left are period-plausible reconstructions, not recovered detail." Faces "can read slightly beautified on very degraded sources". Colour "is inferred from content, or from a reference image", and it is out of scope to recover a creative grade from monochrome. Every tile receives the whole prompt, so a named object "can be painted into a tile that does not contain it"; with modern objects in the negative prompt, empty regions come back as period-plausible background instead of shipping containers.

The replies to the repost were less polite: "It alters the reality of the original image. It's not a restoration." Several pointed at a fork in the boy's hand that they read as a spoon in the scan. I cannot settle that from a 640-pixel crop, which is the point.

So the defensible uses are the ones the card lists: delivering old footage at HD to 4K for a documentary or a remaster, where the audience knows it is looking at a treatment. For archival work it is a rendering, not a record. Keep the scan, ship the two side by side, and caption the colour as inferred, as the card does. When a likeness or a sign matters, give the model a real photograph at the −1 token rather than trusting it to guess.

The relight LoRA: a thin card, a talkative header

RunningHub's card is a few lines: a Chinese title, 光影溶图 ("light-and-shadow blending"), a trigger word and one prompt. The prompt is the instruction the model was trained on:

pengyu  Apply consistent lighting to the product or object, enhance specular highlights,front-back spatial relationship,add natural contact shadow, and blend it seamlessly into the background。

The header's metadata says more than the card. It names the author's ModelScope repository, the source "Modelscope X MuseAI", base model Qwen/Qwen-Image-2.1, lora_rank 32, learning_rate 3e-05, training_steps 2000, max_pixels 1048576 and remove_prefix_in_ckpt pipe.dit. (measured). Those last two are the arguments of DiffSynth-Studio's Qwen-Image-2.1 training script, whose commented edit recipe feeds an edit_image beside each target (measured from the repository). So this is an edit LoRA in the usual Qwen-Image-2.1 sense: the composite goes in as the reference image, the relit picture is the target, at up to 1 MP, and at 0.3 times the script's default learning rate of 1e-4.

What the training pairs were, the card does not say. A natural recipe is to cut a product out of a real photograph, flatten or alter its lighting, paste it back, and ask for the original; the demos' white sticker borders look like that kind of input. That is a guess, not something any file states.

The demos, from the repost's clip, show the claim and its edges.

Left: a yellow plush star toy with pink boots, cut out with a ragged white sticker border, pasted over a neon-lit cyberpunk interior. Right: the same toy without the border, sitting on the machine, lit pink and blue from the neon, with a shadow under its boots.
A plush star pasted with its cut-out border, then relit: the border is gone and the neon colours reach the fur (demo clip in the repost; the same comparison is in the card's poster.gif).
Left: a water bottle and a green soda bottle, flat and evenly lit, pasted on a gold tablecloth in front of a red wall with window shadows. Right: both bottles transparent with the red wall showing through, highlights and shadows on the cloth; the soda's liquid now reads amber.
Two bottles on a table: they become transparent and cast shadows, and the green soda's liquid comes out amber (demo clip in the repost).
Left: a hair dryer with a pink-to-silver gradient finish and a legible brand name, pasted onto cream fur. Right: the dryer in a uniform matte pink, the brand name blurred to illegible marks, with a soft shadow on the fur.
A hair dryer on fur: a shadow appears, but the pink-to-silver finish becomes uniform pink and the printed brand name is no longer legible (demo clip in the repost).
Left: a woman holding out an opaque-looking white plastic bottle against a blue sky. Right: the same woman holding a clear bottle with sunlight through it, narrower, with her fingers redrawn around it.
A bottle held out to camera: the glare is gone and sunlight passes through it, but the bottle's outline and the hand around it are redrawn (demo clip in the repost).

I measured what can be measured from the clip, six input-output pairs at 640 x 640 after X's re-encoding. Phase correlation finds zero global shift in all six, so the poster's "无偏移" ("no offset") holds for the frame. The local change is large: the share of pixels that moved by more than 30 of 255 levels runs from 9.3% (a serum bottle) to 42.7% (the hair dryer, whose fur background is relit too) (measured). The figures show what those pixels are. The model does not composite light onto the product; it redraws the product. Sometimes that is the point, a white-filled bottle becoming glass. Sometimes it is the product's own finish, a legible logo, or the colour of the drink inside.

For product photography that distinction is the whole job. A reply put it in one line: "Great lighting but changing the geometry of a product is really bad." The practical fix is the one every outpainting workflow already uses: keep the model's lighting and shadows outside the product, and paste the original product back on top. Whether the relit highlights survive that paste depends on the shot.

The licensing is unclear in a way that matters for the same audience. The card says to "follow the original project or upstream license", and the upstream is Qwen-Image-2.1's non-commercial research licence.

Running them

Neither fits this site's CPU box in a useful time, so this piece runs nothing. For reference, the paths are short. The restore card's ComfyUI route is the LTX-2.5_V2V_TiledFusion_Upscale.json workflow from ComfyUI-LTXVideo with the LoRA swapped in at 1.0, tile 960 x 544, output FullHD, and a second pass with the Detail Refine IC-LoRA at tile 1024 x 576 for 4K. The order matters: "refining first sharpens the damage and desaturates the colour". The relight LoRA loads with an ordinary LoRA loader on Qwen-Image-2.1 in ComfyUI, with the trigger prompt above and the composite as the reference image. For the memory side of running a 22B video model at all, see the WanGP piece; for few-step Qwen sampling, the few-step piece.

The take

An IC-LoRA is a small idea with a large reach. Put a clean reference into the generator's own sequence, on the same coordinates as the output, and a low-rank diff teaches the model a mapping from one to the other: degraded to clean, pasted to lit, gray to filled. Lightricks ships more than twenty of these for LTX, from deblur to day-to-night, and RunningHub's file shows the same trick on an image editor, trained in 2,000 steps.

What the trick cannot do is add information. The restore LoRA draws twenty times the pixels it is given and every colour, and its card is candid that the result is a reconstruction. The relight LoRA redraws products along with their light. Both are strong editing tools. Neither is a measurement of what was there. For the block both sit on, see the diffusion transformer explainer, and for how shared rotary coordinates line a reference up with its target, the RoPE explainer.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Restore and relight LoRAs: teaching a generator to start from your picture", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026restoreandrelightloras,
  author = {Satyajit Ghana},
  title  = {Restore and relight LoRAs: teaching a generator to start from your picture},
  url    = {https://ai.thesatyajit.com/articles/restore-and-relight-loras},
  year   = {2026}
}
share