Phase 1 gave images a warp grid. Video could not follow, because a
QGraphicsVideoItem cannot be cut into cells — doing it means taking frames
through a QVideoSink and painting them ourselves.
This is the measurement that decides whether that is worth building, and the
tool that produced it: tools/mesh_video_spike.py. Run it on the machine
that will drive the projector — decode is where hardware differs most.
python tools/mesh_video_spike.py --video yourclip.mp4
python tools/mesh_video_spike.py --video yourclip.mp4 --size 3840x2160
python tools/mesh_video_spike.py --video yourclip.mp4 --fast # smoothing off
A 1080p30 and a 2160p30 H.264 clip, decoded through QVideoSink and drawn as
an N x N warped grid into a full-size buffer. Three routes were timed per
frame, plus decode on its own, plus the path that ships today as a baseline.
The machine: 4-core Xeon at 2.1 GHz, software decode, no RHI backend (Qt reported "Using CPU conversion"). A Windows show laptop with a GPU should do better on conversion and probably on the warp too. Treat these as a floor.
| Route | 1x1 | 4x4 | 16x16 |
|---|---|---|---|
toImage() then draw |
17.4 ms | 20.6 ms | 31.8 ms |
QVideoFrame.paint() straight |
12.0 ms | 15.6 ms | 27.4 ms |
| Budget at 60 fps | 16.7 ms | 16.7 ms | 16.7 ms |
| Budget at 30 fps | 33.3 ms | 33.3 ms | 33.3 ms |
Baseline, the path that ships today, same clip full-canvas:
| What | Per rendered frame |
|---|---|
| Plain playback, no warp | 1.3 ms |
| Corner pinned (keystone) | 5.5 ms |
It is viable at 1080p. One mesh video layer fits inside a 30 fps budget at every subdivision, and inside 60 fps up to about 4x4.
Subdivision is not the expense. Going from one cell to 16x16 barely more than doubles the cost. The expense is resampling a 1080p frame projectively at all: a single warped cell already costs 12 ms, where the same frame unwarped costs 1.3 ms. A fine grid is nearly free once you are paying for the warp.
QVideoFrame.paint() is the route. It converts internally and saves the
whole 5.1 ms toImage() step for no cost in draw time. Anything built here
should use it rather than converting by hand.
Smoothing is the biggest lever. Turning it off cuts the draw roughly in half — 16x16 drops from 26 ms to 11 ms. That is a quality decision rather than a free win: nearest-neighbour sampling on a warped projection aliases visibly on detailed content. Worth exposing per object, not worth defaulting to.
It costs about twice what a keystone costs today — 12 ms against 5.5 ms. That is the honest price of the feature, and it buys surfaces a corner pin cannot fit at all.
4K does not fit. 60 ms a frame at one cell, 100 ms at 16x16, or 10-17 fps. Not on this machine, and probably not on any CPU raster path. 4K mesh video needs the GPU, which is a different piece of work.
Both were expected to matter, and neither did. Recorded because the next person to look at this will have the same two ideas.
Giving each cell only its own slice of the frame made no difference — 11.82 ms against 11.86 ms. The theory was that clipping still hands the whole frame to the transform, so N cells transform it N times. Qt's clip evidently already avoids that work. It also means the Phase 1 image renderer, which draws the whole buffer per cell and clips, is not leaving anything on the table.
Wrapping the frame's own bits without converting is not available. Frames
arrive as Format_YUV420P, which no QImage format can wrap. The measurement
that suggested it cost 0.05 ms was wrapping YUV bytes in an RGB image and
producing garbage quickly. The tool now checks the format and reports "n/a"
rather than a number that means nothing.
Build it, for 1080p, using QVideoFrame.paint(), with smoothing exposed per
object — but run the tool on the show laptop first. The number that decides it
is the one from the machine that has to hold it for two hours, not this one.
Phase 3 shipped on this basis. Three things the spike did not catch, recorded so the numbers above are read with them in mind.
The frame has to be held, not read back. QVideoSink.videoFrame() is only
dependable while videoFrameChanged is being delivered; a moment later the
backend may have recycled the buffer, and painting from it draws nothing.
MeshVideoLayerItem keeps a reference to the frame it was handed. That is also
why it lets go of it on unload — an off-air scene must not pin a full-size
buffer.
QVideoFrame.paint() does honour the painter's clip and transform. It was
briefly believed not to, on the strength of a render that looked like the warp
had come apart. It had not: the test clip was ffmpeg's testsrc2, which
contains a magenta bar, and the magenta background being used to find gaps was
indistinguishable from it. Twice. Solid-color clips, or the two-background
coverage trick in tests/test_mesh_video.py, avoid the trap.
Composing into a buffer first costs about 6.5 ms a frame at 1080p, measured through the real layer rather than the spike's harness. So the shipped renderer draws each cell straight from the frame, and only composes when it must: an aiming grid has to be warped with the picture to be any use, and a vignette applied after the warp would be the wrong shape. Both pay the 6.5 ms, and neither is on during a show.
Through the real layer the per-cell route measures 12.2 ms at one cell, 15.2 at 4x4 and 21.5 at 8x8 — close to the spike. 16x16 comes out at 45.8 ms rather than 27.4, because the spike drew into a bare canvas with no scene graph around it. Keep 1080p mesh video at 8x8 or below on hardware like this.