# Restore and relight LoRAs: teaching a generator to start from your picture

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/restore-and-relight-loras
> date: 2026-10-06
> tags: video-generation, image-generation, diffusion, lora, fine-tuning, open-weights, explainer

Two LoRAs came out in the last week of September, and on the surface they have nothing in common.
Lightricks' [LTX-2.5 Restore IC-LoRA](https://huggingface.co/Lightricks/LTX-2.5-22b-IC-LoRA-Restore)
takes archive video, "low-resolution web rips, low-bitrate broadcast transfers, tape, sepia or
black-and-white film scans", and renders the same shot clean and in colour. RunningHub's
[Qwen-Image-2.1 lighting LoRA](https://huggingface.co/RunningHubAI/rh-qwen-image-2.1-lora-2104918997757157378)
takes a product pasted onto a background and, in the repost's words, removes reflections and glare,
enhances specular highlights and adds natural shadows.

Underneath they are the same move. A model trained to generate from noise and text is shown a
picture, clean, inside its own token sequence, and a small set of extra weights teaches it to treat
that picture as the thing to redraw. This piece explains that move from the matrix up, then checks
each release against what can be read without a GPU: the cards, the safetensors headers over HTTP
range requests, the LTX-2 trainer and pipeline code, DiffSynth-Studio, and ComfyUI's loader.

One constraint shaped the work. The Lightricks repository is gated behind a click-through, and this
box has no Hugging Face token, so its weights, its README and its example videos all return 401. The
card text is public on the model page, the file sizes are public in the API, and the comparison clip
is in the repost. Everything about the restore weights below is therefore reasoned from a byte count,
and labelled so.

<ModelCard
  repo="Lightricks/LTX-2.5-22b-IC-LoRA-Restore"
  claimed="Restores and colourises archive footage on LTX-2.5; PSNR +2.0 dB over the degraded input on 8 validation pairs"
  note="Gated; weights not read. The file is 1,711,643,914 bytes (measured from the Hub API), which matches rank 128 on all six attention modules of all 48 blocks, 855,638,016 parameters, to within 9,127 bytes of header (reasoned). Trained on LTX-2.3, tested on LTX-2.5 unchanged (reported). The validation sources also appear in training under other degradations (reported). LTX-2 Community License."
/>

<ModelCard
  repo="RunningHubAI/rh-qwen-image-2.1-lora-2104918997757157378"
  claimed="Consistent lighting for a pasted object: no reflections or glare, specular highlights, natural shadows"
  note="Measured from the header: 448 BF16 tensors, rank 32 on seven linear layers in each of 32 blocks, 83,886,080 parameters, 1.18% of the 7,115,124,736-parameter denoiser. Metadata: trained on ModelScope with DiffSynth-Studio settings, learning rate 3e-05, 2,000 steps, trigger word pengyu. The card describes no data, metrics or licence of its own; the base model's is the non-commercial Qwen Research License."
/>

## A LoRA is a low-rank diff

Fine-tuning a layer means changing its weight matrix $W$, of size $d_{out} \times d_{in}$. LoRA
([Hu et al., 2021](https://arxiv.org/abs/2106.09685)) freezes $W$ and learns the change as a product
of two thin matrices:

$$
W' = W + \lambda \cdot \frac{\alpha}{r} \, B A, \qquad A \in \mathbb{R}^{r \times d_{in}},\; B \in \mathbb{R}^{d_{out} \times r}
$$

$r$ is the rank, $\alpha$ a fixed scale chosen at training time, and $\lambda$ the strength slider
in ComfyUI. $B$ starts at zero, so the adapter starts as a no-op. The cost is
$r(d_{in} + d_{out})$ parameters per layer instead of $d_{in} d_{out}$. For a 4,096 by 4,096
projection at rank 128 that is 1,048,576 numbers against 16,777,216, 6.25% of the layer (reasoned).
At rank 32 it is 262,144, or 1.56%.

The bet is that the change a task needs lies in a few directions per layer. Restoration and
relighting are good candidates: the base model already knows what a steam locomotive or a glass
bottle looks like under daylight. The adapter has to teach it something narrower: where to look for
the layout, and what to do differently once it has.

Both files set $\alpha = r$, so $\alpha / r = 1$ and the strength slider is the only scale. The
restore card states rank 128, alpha 128. The relight file has no alpha tensor, and DiffSynth-Studio's
`add_lora_to_model` defaults `lora_alpha` to the rank; ComfyUI applies a scale of 1 to a file with no
alpha (all measured from the sources).

## What is in the two files

The relight header is 58,240 bytes of JSON, read with two range requests. It lists 448 BF16 tensors:
an `A` and a `B` for `to_q`, `to_k`, `to_v`, `to_out.0`, `img_mlp.gate_layer`, `img_mlp.proj` and
`img_mlp.out` in each of 32 blocks. Each attention projection is 4,096 square; the two MLP inputs go
4,096 to 12,288 and the output comes back. That is 2,621,440 parameters a block and 83,886,080 in
all, and 8 + 58,240 + 2 x 83,886,080 is the file's size to the byte, 167,830,408 (measured).

The restore header cannot be read. What can be read is the base model. The LTX-2.3 transformer header
lists 21,005,004,544 parameters and six attention modules in each of 48 blocks: video self-attention,
video-to-text cross-attention, audio self-attention, audio-to-text cross-attention, and the two
cross-modal modules, audio-to-video and video-to-audio (measured). The LTX-2 paper's own diagram shows
them:

<Figure
  src="https://ai.thesatyajit.com/articles/restore-and-relight-loras/fig1.png"
  alt="Two diagrams. On the left, one LTX-2 block as two parallel stacks: the video stack runs self attention, T2V cross attention, A2V cross attention and an FFN on the video hidden state; the audio stack runs self attention, T2A cross attention, V2A cross attention and an FFN on the audio hidden state; text hidden state feeds both cross attentions. On the right, a detailed view of the bidirectional audio-video cross attention, with Q, K and V projections, temporal 1D RoPE, scale and shift from the timesteps, and a gate."
  caption="The dual-stream LTX-2 block: video and audio each have self-attention, text cross-attention and a cross-modal attention, six attention modules in all, every one with to_q, to_k, to_v and to_out (LTX-2 paper, Figure 2)."
/>

The card says the restore LoRA "targets to_q/to_k/to_v/to_out". In the LTX-2 trainer that is a
suffix match, and the trainer's configuration guide warns that the short patterns "will match all
attention modules including `attn1.to_k`, `audio_attn1.to_k`, `audio_to_video_attn.to_k`, and
`video_to_audio_attn.to_k`" (measured from the docs). At rank 128 over all six modules of all 48
blocks, that is 855,638,016 parameters, 1,711,276,032 bytes of BF16. Rebuild the header those
tensors would need, with the licence text the ungated Union-Control LoRA carries as metadata, and the
predicted file comes to 1,711,634,787 bytes. The real file is 1,711,643,914: 9,127 bytes apart, less
than one part in 180,000 (reasoned). Video attention alone would be 805,306,368 bytes, half the file,
so the short patterns are what was used.

That has a consequence the card does not mention. The restore run trained on video; the card lists
"Audio: Not trained for audio generation". With no audio in the batch, the LTX-2 transformer skips
the audio and cross-modal branches entirely (`run_ax`, `run_a2v` and `run_v2a` are false), so their
LoRA `B` matrices get no gradient and keep their zero initialisation. If that is what happened, 452,984,832
of the 855,638,016 parameters, 53%, are zeros that load, cost memory, and change nothing
(reasoned from the trainer code; not checked against the weights).

The calculator below rebuilds every one of these files from the measured layer shapes. Pick a base,
tick the layer groups, slide the rank. It reproduces the two ungated LTX-2.3 IC-LoRA headers I read
(Lightricks' Union-Control at rank 64, 327,155,712 parameters, and a community colourizer at rank 32,
163,577,856, both on video attention and the video feed-forward) and both Qwen-Image-2.1 LoRAs on this
site exactly.

<LoraBudget />

On Qwen the same layer budget can be spent two ways. Qwen-Image-2.1's MLP input is one fused
`gate_up` weight in ComfyUI and ai-toolkit, but two separate 12,288-wide layers in diffusers and
DiffSynth-Studio. [ausboss's outpaint LoRA](/articles/qwen-image-2-1-outpaint-lora) targets the fused
one: one rank-32 adapter, 79,691,776 parameters. The relight LoRA targets the halves: two independent
rank-32 adapters, 4,194,304 parameters more (measured). ComfyUI's loader has a branch for exactly
this, mapping `gate_layer` and `proj` onto the two halves of the fused weight, and accepts PEFT's
`lora_B.default.weight` naming, so the file loads as shipped (measured from `comfy/lora.py` and
`weight_adapter/lora.py` at commit `d49e888`).

## In context: the reference rides in the sequence

A LoRA changes weights. It does not, by itself, give a text-to-video model a way to see an input
video. That is what the "IC" adds. The name comes from Alibaba's
[In-Context LoRA](https://arxiv.org/abs/2410.23775), which tiled related images into one canvas;
Lightricks' version is more direct, and the LTX-2 trainer spells it out in
`training_strategies/video_to_video.py` (measured from the code):

1. Encode the target clip and the reference clip with the same video VAE. The VAE compresses 32x in
   height and width and 8x in time, and the patch size is 1, so a 960 x 544 x 97 window is
   30 x 17 x 13 = 6,630 tokens.
2. Noise only the target. The reference stays clean, and its per-token timestep is 0.
3. Concatenate the two into one sequence and give both the same position grid. At downscale factor
   1, reference token (t, h, w) and target token (t, h, w) carry the same rotary coordinates.
4. Run the transformer over the lot, with ordinary bidirectional self-attention.
5. Compute the flow-matching loss on the target tokens only.

At inference `VideoConditionByReferenceLatent` does the same thing: it appends the reference tokens
with a denoise mask of 1 − strength, so a guide strength of 1 keeps them clean at every step. The
order differs from training, reference first there and appended here, which attention does not care
about: it sees positions, not order (reasoned).

That is the whole mechanism. There is no encoder for the reference and no new input channel. The
target attends to the reference in every layer of the same network, and because the two share
coordinates, "the matching pixel in the archive clip" is one rotary rotation away from every target
token. What the LoRA learns is how to use that path: which parts of the reference to copy (layout,
motion, structure) and which to override (grain, cast, blur, the absence of colour).

<IcSequence />

The cost is plain in the token count. At downscale factor 1 the reference is as long as the target,
so every linear layer processes twice the tokens and self-attention four times the pairs: 13,260
tokens for one 960 x 544 tile of 97 frames (reasoned). Lightricks offers a cheaper variant elsewhere:
the Union-Control LoRA's metadata says `reference_downscale_factor` 2, a quarter of the reference
tokens, with positions scaled back up to the target's grid. The restore LoRA's card says factor 1,
which makes sense for a task whose whole point is detail at the target's scale.

There is also a dial between full and no attention. Wrap a reference in
`ConditioningItemAttentionStrengthWrapper` with strength `s` below 1, and `build_attention_mask`
writes `s` into the target-to-reference blocks, 1 inside each reference, and 0 between two different
references. The float mask becomes an additive bias of $\log s$ on the attention logits, so `s` = 0.5
halves the unnormalised weight of every target-reference pair (measured from `mask_utils.py` and
`transformer_args.py`). At `s` = 1 no mask is built at all.

Qwen-Image-2.1 edits the same way with a different mask. Its reference image is a prefix in one
stream of 32 blocks, modulated from `t = 0`, and the mask is block-causal: the target reads the
reference, the reference never reads the target, as
[the release piece](/articles/qwen-image-2-1#block-causal-attention-and-why-the-cache-is-exact) covers.

<Figure
  src="https://ai.thesatyajit.com/articles/qwen-image-2-1/fig1.png"
  alt="An attention mask drawn as a grid of query rows against key columns, divided into four segments: system prefix, optional input image, edit instruction, target image. The text segments are causal staircases; the image segments are solid blocks, bidirectional inside themselves. The target image rows see every column."
  caption="Qwen-Image-2.1's mixed-granularity mask. For the relight LoRA, the pasted composite is the input image and the relit picture is the target (Qwen-Image-2.1 release, architecture figure)."
/>

So the relight LoRA is in-context too, in everything but name. The difference that matters is in the
arrow: LTX's reference and target see each other, Qwen's reference is read-only, which is what lets
Qwen cache it.

## How the restore LoRA was trained

The card is unusually complete about this, and it is the part to read before using it.

**Synthetic pairs.** "A clean 1080p clip was turned monochrome or tinted (neutral broadcast, cool
tape, warm sepia), blurred as a period lens would, downscaled to 240p to 360p, encoded at a low
bitrate one to three times over, sometimes denoised, scaled back up, and given film damage." The
degraded clip is the reference and the clean clip the target. There are 151 training pairs and 8
held out, 97 frames at 24 fps, one shared caption (reported).

**The recipe.** The LTX-2 trainer's flexible strategy: the degraded video always present as the
reference, the first frame given clean with probability 0.25, and the first 17 frames given as a
prefix with probability 0.5. Rank 128, alpha 128, the Prodigy optimiser at learning rate 1.0, a
cosine schedule, batch 1, 4,000 steps, buckets of 960 x 544 at 97 and 49 frames (reported). The
prefix training is what lets a long clip be restored as chained 97-frame windows, each starting from
the previous window's last 17 restored frames.

**The settings it wants.** Eight steps on the distilled model, guidance 1.0, LoRA strength 1.0, and
the clip upscaled with lanczos to the working canvas before it becomes the guide, so "the model never
sees the small scan". The canvas is cut into overlapping 960 x 544 tiles, the trained size, fused at
every step, even for a full-HD output. The distilled sigmas are 1.0, 0.99375, 0.9875, 0.98125, 0.975,
0.909375, 0.725, 0.421875 and 0.0, so five of the eight steps run above 0.9 (reasoned): most of the
schedule is spent where layout and colour are decided. Strength "behaves as a switch: at 0.7 the
restoration turns off and the clip stays monochrome; 1.4 over-tints" (reported).

The card's gallery is three public-domain clips: a 1954 family meal, the Lumière train and the
Hindenburg. Two of them, as they appeared in the repost:

<Figure
  src="https://ai.thesatyajit.com/articles/restore-and-relight-loras/restore-meal.jpg"
  alt="A side-by-side crop. Left, labelled SOURCE 640x480, upscaled: a soft, grainy black-and-white frame of a boy in a dark sweater leaning over a table with a fork, a girl's hair in the foreground. Right, labelled RESULT 2880x2176, 1:1: the same pose in colour and sharp, a dark teal sweater, an orange wall, skin tones, individual strands of hair, a fork with food on it."
  caption="The 1954 family meal: the scan lanczos-upscaled on the left, the restored and refined result on the right, a 1:1 crop of the 2880 x 2176 output. The wall's orange, the sweater's teal and every strand of hair are the model's (Lightricks Restore IC-LoRA model card, gallery clip)."
/>

<Figure
  src="https://ai.thesatyajit.com/articles/restore-and-relight-loras/restore-hindenburg.jpg"
  alt="A side-by-side crop. Left, labelled SOURCE: the grey ribbed nose of the Hindenburg against a grey sky, with blurred trees and a mast below. Right, labelled RESULT: the same airship in blue-silver with crisp panel ribs, a blue sky, green trees and a yellow lattice mooring mast."
  caption="The Hindenburg at its mooring mast in 1937: the same crop before and after. The yellow mast is a colour decision, made from content and the prompt (Lightricks Restore IC-LoRA model card, gallery clip)."
/>

**The numbers.** On 8 validation pairs at 1920 x 1056, 97 frames, stage 1 only: PSNR +2.0 dB over
the degraded input, a Lab colour error from 24 to 16, and high-frequency energy from 61% to 85% of the
target, against +0.8 dB and 75% for the previous version (reported). Read the fine print next to
them: those 8 pairs' "sources also appear in training under other degradations". The model has seen
those clean frames. On clips whose sources it never saw, the card says, it "does not beat the aligned
input on pixel PSNR; judge those by eye". That is the honest version of the claim, and it is
Lightricks' own.

A reference image helps where it exists. Real crops of a sign, a logo or a face, placed over their
own region on a grey canvas and attached at frame index −1, raised the sign region by 4.6 dB and the
faces by 4.0 dB against the truth on one street clip (reported). It is the one mechanism here that
adds information instead of inferring it.

## What "restore" means here

Count what the model is given. The source is 640 x 480, one channel. The output is 2880 x 2176, three
channels. That is 20.4 times the pixels, and two of the three colour channels have no source at all
(reasoned). Everything between the scan's samples is drawn by a video generator conditioned on them.

The card says so in its own words. On the train clip, "the station side of the frame holds almost no
information in the scan, so the model rebuilds it: the buildings, the hillside and the porter's cart
on the left are period-plausible reconstructions, not recovered detail." Faces "can read slightly
beautified on very degraded sources". Colour "is inferred from content, or from a reference image",
and it is out of scope to recover a creative grade from monochrome. Every tile receives the whole
prompt, so a named object "can be painted into a tile that does not contain it"; with modern objects
in the negative prompt, empty regions come back as period-plausible background instead of shipping
containers.

The replies to the repost were less polite: "It alters the reality of the original image. It's not a
restoration." Several pointed at a fork in the boy's hand that they read as a spoon in the scan. I
cannot settle that from a 640-pixel crop, which is the point.

So the defensible uses are the ones the card lists: delivering old footage at HD to 4K for a
documentary or a remaster, where the audience knows it is looking at a treatment. For archival work it
is a rendering, not a record. Keep the scan, ship the two side by side, and caption the colour as
inferred, as the card does. When a likeness or a sign matters, give the model a real photograph at
the −1 token rather than trusting it to guess.

## The relight LoRA: a thin card, a talkative header

RunningHub's card is a few lines: a Chinese title, 光影溶图 ("light-and-shadow blending"), a trigger
word and one prompt. The prompt is the instruction the model was trained on:

```text
pengyu  Apply consistent lighting to the product or object, enhance specular highlights，front-back spatial relationship，add natural contact shadow, and blend it seamlessly into the background。
```

The header's metadata says more than the card. It names the author's ModelScope repository, the
source "Modelscope X MuseAI", base model `Qwen/Qwen-Image-2.1`, `lora_rank` 32, `learning_rate`
3e-05, `training_steps` 2000, `max_pixels` 1048576 and `remove_prefix_in_ckpt` `pipe.dit.`
(measured). Those last two are the arguments of DiffSynth-Studio's Qwen-Image-2.1 training script,
whose commented edit recipe feeds an `edit_image` beside each target (measured from the repository).
So this is an edit LoRA in the usual Qwen-Image-2.1 sense: the composite goes in as the reference
image, the relit picture is the target, at up to 1 MP, and at 0.3 times the script's default learning rate
of 1e-4.

What the training pairs were, the card does not say. A natural recipe is to cut a product out of a
real photograph, flatten or alter its lighting, paste it back, and ask for the original; the demos'
white sticker borders look like that kind of input. That is a guess, not something any file states.

The demos, from the repost's clip, show the claim and its edges.

<div className="grid gap-x-4 sm:grid-cols-2">

<Figure
  src="https://ai.thesatyajit.com/articles/restore-and-relight-loras/relight-star.jpg"
  alt="Left: a yellow plush star toy with pink boots, cut out with a ragged white sticker border, pasted over a neon-lit cyberpunk interior. Right: the same toy without the border, sitting on the machine, lit pink and blue from the neon, with a shadow under its boots."
  caption="A plush star pasted with its cut-out border, then relit: the border is gone and the neon colours reach the fur (demo clip in the repost; the same comparison is in the card's poster.gif)."
/>

<Figure
  src="https://ai.thesatyajit.com/articles/restore-and-relight-loras/relight-shelf.jpg"
  alt="Left: a water bottle and a green soda bottle, flat and evenly lit, pasted on a gold tablecloth in front of a red wall with window shadows. Right: both bottles transparent with the red wall showing through, highlights and shadows on the cloth; the soda's liquid now reads amber."
  caption="Two bottles on a table: they become transparent and cast shadows, and the green soda's liquid comes out amber (demo clip in the repost)."
/>

<Figure
  src="https://ai.thesatyajit.com/articles/restore-and-relight-loras/relight-dryer.jpg"
  alt="Left: a hair dryer with a pink-to-silver gradient finish and a legible brand name, pasted onto cream fur. Right: the dryer in a uniform matte pink, the brand name blurred to illegible marks, with a soft shadow on the fur."
  caption="A hair dryer on fur: a shadow appears, but the pink-to-silver finish becomes uniform pink and the printed brand name is no longer legible (demo clip in the repost)."
/>

<Figure
  src="https://ai.thesatyajit.com/articles/restore-and-relight-loras/relight-bottle.jpg"
  alt="Left: a woman holding out an opaque-looking white plastic bottle against a blue sky. Right: the same woman holding a clear bottle with sunlight through it, narrower, with her fingers redrawn around it."
  caption="A bottle held out to camera: the glare is gone and sunlight passes through it, but the bottle's outline and the hand around it are redrawn (demo clip in the repost)."
/>

</div>

I measured what can be measured from the clip, six input-output pairs at 640 x 640 after X's
re-encoding. Phase correlation finds zero global shift in all six, so the poster's "无偏移" ("no
offset") holds for the frame. The local change is large: the share of pixels that moved by more than
30 of 255 levels runs from 9.3% (a serum bottle) to 42.7% (the hair dryer, whose fur background is
relit too) (measured). The figures show what those pixels are. The model does not composite light
onto the product; it redraws the product. Sometimes that is the point, a white-filled bottle becoming
glass. Sometimes it is the product's own finish, a legible logo, or the colour of the drink inside.

For product photography that distinction is the whole job. A reply put it in one line: "Great
lighting but changing the geometry of a product is really bad." The practical fix is the one every
outpainting workflow already uses: keep the model's lighting and shadows outside the product, and
paste the original product back on top. Whether the relit highlights survive that paste depends on
the shot.

The licensing is unclear in a way that matters for the same audience. The card says to "follow the
original project or upstream license", and the upstream is Qwen-Image-2.1's
[non-commercial research licence](/articles/qwen-image-2-1#the-licence-is-the-news).

## Running them

Neither fits this site's CPU box in a useful time, so this piece runs nothing. For reference, the
paths are short. The restore card's ComfyUI route is the `LTX-2.5_V2V_TiledFusion_Upscale.json`
workflow from ComfyUI-LTXVideo with the LoRA swapped in at 1.0, tile 960 x 544, output FullHD, and a
second pass with the Detail Refine IC-LoRA at tile 1024 x 576 for 4K. The order matters: "refining
first sharpens the damage and desaturates the colour". The relight LoRA loads with an ordinary LoRA
loader on Qwen-Image-2.1 in ComfyUI, with the trigger prompt above and the composite as the reference
image. For the memory side of running a 22B video model at all, see
[the WanGP piece](/articles/vram-is-a-policy); for few-step Qwen sampling,
[the few-step piece](/articles/qwen-image-2-1-few-step).

## The take

An IC-LoRA is a small idea with a large reach. Put a clean reference into the generator's own
sequence, on the same coordinates as the output, and a low-rank diff teaches the model a mapping from
one to the other: degraded to clean, pasted to lit, gray to filled. Lightricks ships more than twenty
of these for LTX, from deblur to day-to-night, and RunningHub's file shows the same trick on an image
editor, trained in 2,000 steps.

What the trick cannot do is add information. The restore LoRA draws twenty times the pixels it is
given and every colour, and its card is candid that the result is a reconstruction. The relight LoRA
redraws products along with their light. Both are strong editing tools. Neither is a measurement of
what was there. For the block both sit on, see
[the diffusion transformer explainer](/architectures/diffusion-transformer), and for how shared
rotary coordinates line a reference up with its target, [the RoPE explainer](/architectures/rope).
