Moiré Video Authentication: A Physical Signature Against AI Video Generation

ECCV 2026

*Equal contribution

Boston University
The Moiré effect emerges from two overlaid periodic layers; fringes in real video shift predictably with camera motion, while AI-generated video produces distorted, non-physical fringe shifts.

A compact two-layer grating produces Moiré fringes whose motion is governed by deterministic optics. Real cameras reproduce this coupling exactly; state-of-the-art video generators do not.

Abstract

Recent advances in video generation have made AI-synthesized content increasingly difficult to distinguish from real footage. We propose a physics-based authentication signature that real cameras produce naturally, but that generative models cannot faithfully reproduce. Our approach exploits the Moiré effect: interference fringes formed when a camera views a compact two-layer grating structure. We derive the Moiré motion invariant, showing that fringe phase and grating image displacement are linearly coupled by optical geometry, independent of viewing distance and grating structure. A verifier extracts both signals from video and tests their correlation. We validate the invariant on real-captured, physically-rendered, and AI-generated videos from multiple state-of-the-art generators (Veo 3.1, Grok Imagine, LTX-2), and find significantly different correlation signatures between real and generated content (μ = 0.87 vs. 0.57, p < 10−20).

Demo Video

Introduction

AI-generated video has become realistic enough to cause real-world harm — from a 2024 finance-department fraud in which employees were deceived by a synthetic video call, to a 2025 hoax impersonating a national presidential candidate. As generation quality improves, reliable authenticity verification becomes critical. Existing approaches largely fall into two camps: post-hoc forensic detectors that hunt for pixel-level artifacts or learned fingerprints, and digital watermarking that embeds metadata (e.g., SynthID). Both are fragile — artifacts vanish as models improve, and watermarks are lost under re-encoding, cropping, or redistribution.

We take a different, physics-grounded route. A real camera capturing a physical scene obeys the laws of optics; a generative model only learns image statistics. We exploit this gap with a passive Moiré signature: a compact two-layer grating placed in the scene produces interference fringes whose motion is deterministically coupled to camera motion. Unlike active structured-light methods, our approach requires no emitter, projector, or external hardware — only ordinary relative motion between the camera and the assembly. By testing whether fringe-phase variation stays correlated with grating image displacement, we obtain a verifiable, geometry-driven signal that separates authentic capture from synthetic generation.

Passive Moiré-based signature (requiring only camera movement) versus active structured-light methods that require external emitter hardware.

Our passive Moiré signature needs only camera motion, in contrast to active structured-light approaches that require dedicated emitters.

The Moiré Motion Invariant

Two overlaid gratings with periods pf and pr produce beat-period Moiré fringes of period pm = pf pr / |pf − pr|, magnifying tiny relative shifts between the layers into large, easily-measured fringe motion.

When the camera translates by Δxc at distance D from an assembly with layer gap g, the two layers shift relative to one another, changing the fringe phase by Δφ = ±(2π g / pr D) Δxc. The phase change depends on the unknown distance D — but so does the grating's apparent image-plane displacement. Dividing the two cancels D entirely, yielding the correlation invariant for pure translation:

Δφ = ±(2π g / pr f) Δutrans   ⟹   |ρ| = |Corr(Δφ, Δutrans)| ≈ 1.

Here f is the focal length and Δutrans the translational image displacement. Crucially, the relationship is independent of viewing distance and grating structure: any authentic recording exhibits near-perfect correlation between fringe phase and translational displacement, while pure rotation induces no inter-layer parallax and hence no fringe shift. A generator cannot reproduce this coupling without effectively solving the underlying wave-optics — so synthetic video breaks the invariant.

Proof-of-Concept Implementation

We build the assembly from off-the-shelf parts: a 1 mm lenticular sheet (50 LPI, pf = 0.508 mm) over a printed grating (pr = 0.52 mm), giving a ∼42× Moiré magnification, mounted on an acrylic base with four ArUco markers for pose estimation.

Given an input video, the verifier runs three stages:

  1. Grating tracking & fringe extraction. Detect the ArUco markers, compute a homography to a canonical frame, enhance contrast (CLAHE), and collapse the 2D image into a 1D profile along the fringe direction.
  2. Phase tracking. Extract phase via FFT at the Moiré beat frequency and track it incrementally to avoid 2π wrap-around, accumulating a smooth cumulative phase Φ(t).
  3. Displacement & correlation. Recover camera pose with PnP on the markers, decompose the trajectory into translation and rotation, project translation onto the fringe-sensitive axis, and compute the Pearson correlation over sliding 30-frame windows.
Exploded view of the three-layer grating assembly, ArUco-marker detection isolating the Moiré region, and extracted fringes after canonical transformation.

Assembly, ArUco-based region isolation, and the extracted fringes used for phase tracking.

Main Contributions

  • A novel Moiré-based video authentication framework with a derived, distance-invariant Moiré motion invariant.
  • A practical proof-of-concept built entirely from off-the-shelf materials (lenticular sheet, printed grating, ArUco markers).
  • A three-stage verification pipeline and a threat-model analysis covering splicing and face-swap attacks.
  • Empirical validation across real, physically-rendered, and AI-generated video from multiple state-of-the-art generators.

Experiments & Results

We evaluate on three sources: 87 real clips (iPhone 15, 1080p/60fps) across 12 subjects and 29 scenes with three motion paradigms; 70 physics renderings from Blender Cycles at varying distances (plus 10 pure-translation sequences); and 92 curated AI-generated videos from Veo 3.1, Grok Imagine, and LTX-2. Text-to-video generation failed outright — all 60 attempts were unable to synthesize a recognizable Moiré pattern — so we used a deliberately generous image-to-video protocol (first-frame conditioning, manual region selection, human-corrected tracking, cherry-picked outputs) to give the generators the best possible chance.

The invariant separates real from synthetic cleanly. Real recordings correlate at μ = 0.87 (σ = 0.14) and physics renderings at μ = 0.99, while AI-generated video reaches only μ = 0.57 (σ = 0.21). The gap is highly significant — Welch's t(160) = 11.6, p < 10−20, Cohen's d = 1.71 — and all three generators fail comparably (Grok 0.56, LTX-2 0.57, Veo 0.58).

Distribution of Pearson correlation coefficients across real recordings (mu=0.87), physics-based renderings (mu=0.99), and AI-generated videos (mu=0.57).

Correlation-coefficient distributions: real (μ = 0.87) and physics renderings (μ = 0.99) cluster near 1, while AI-generated video (μ = 0.57) does not.

When Generators Fail

Because generators reproduce appearance rather than optics, they cannot keep Moiré fringes physically consistent over time. Text-to-video outputs collapse immediately — fringes appear as unnaturally thick stripes with severe deformation:

Text-to-video failure (Grok).
Text-to-video failure (Veo).
Text-to-video failure (LTX-2).

Even under the far more generous image-to-video setting, temporal inconsistency, warping, and structural deformation of the grating region persist across frames:

Image-to-video failure (Grok).
Image-to-video failure (Veo).
Image-to-video failure (Veo).

Threat Model & Limitations

We analyze two attacks. Splicing a genuine Moiré region into a fake video fails, because the transplanted fringe phase no longer matches the target's camera motion, breaking the correlation. Face-swaps outside the Moiré region are a genuine limitation, addressed by pairing our signature with complementary deepfake detectors and by proposed protocols such as a "Moiré ID" or spatial-overlap verification. The approach assumes generators do not ray-trace grating optics and that some relative camera–assembly motion is present to generate the signal.

BibTeX

@misc{qing2026moire,
      title={Moir\'e Video Authentication: A Physical Signature Against AI Video Generation},
      author={Yuan Qing and Kunyu Zheng and Lingxiao Li and Boqing Gong and Chang Xiao},
      year={2026},
      eprint={2604.01654},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2604.01654},
}