SurfaceCast

Mesh video: the Phase 2 spike

Phase 1 gave images a warp grid. Video could not follow, because a QGraphicsVideoItem cannot be cut into cells — doing it means taking frames through a QVideoSink and painting them ourselves.

This is the measurement that decides whether that is worth building, and the tool that produced it: tools/mesh_video_spike.py. Run it on the machine that will drive the projector — decode is where hardware differs most.

python tools/mesh_video_spike.py --video yourclip.mp4
python tools/mesh_video_spike.py --video yourclip.mp4 --size 3840x2160
python tools/mesh_video_spike.py --video yourclip.mp4 --fast   # smoothing off

What was measured

A 1080p30 and a 2160p30 H.264 clip, decoded through QVideoSink and drawn as an N x N warped grid into a full-size buffer. Three routes were timed per frame, plus decode on its own, plus the path that ships today as a baseline.

The machine: 4-core Xeon at 2.1 GHz, software decode, no RHI backend (Qt reported "Using CPU conversion"). A Windows show laptop with a GPU should do better on conversion and probably on the warp too. Treat these as a floor.

The numbers, 1080p into 1920x1080, smoothing on

Route 1x1 4x4 16x16
toImage() then draw 17.4 ms 20.6 ms 31.8 ms
QVideoFrame.paint() straight 12.0 ms 15.6 ms 27.4 ms
Budget at 60 fps 16.7 ms 16.7 ms 16.7 ms
Budget at 30 fps 33.3 ms 33.3 ms 33.3 ms

Baseline, the path that ships today, same clip full-canvas:

What Per rendered frame
Plain playback, no warp 1.3 ms
Corner pinned (keystone) 5.5 ms

What it says

It is viable at 1080p. One mesh video layer fits inside a 30 fps budget at every subdivision, and inside 60 fps up to about 4x4.

Subdivision is not the expense. Going from one cell to 16x16 barely more than doubles the cost. The expense is resampling a 1080p frame projectively at all: a single warped cell already costs 12 ms, where the same frame unwarped costs 1.3 ms. A fine grid is nearly free once you are paying for the warp.

QVideoFrame.paint() is the route. It converts internally and saves the whole 5.1 ms toImage() step for no cost in draw time. Anything built here should use it rather than converting by hand.

Smoothing is the biggest lever. Turning it off cuts the draw roughly in half — 16x16 drops from 26 ms to 11 ms. That is a quality decision rather than a free win: nearest-neighbour sampling on a warped projection aliases visibly on detailed content. Worth exposing per object, not worth defaulting to.

It costs about twice what a keystone costs today — 12 ms against 5.5 ms. That is the honest price of the feature, and it buys surfaces a corner pin cannot fit at all.

4K does not fit. 60 ms a frame at one cell, 100 ms at 16x16, or 10-17 fps. Not on this machine, and probably not on any CPU raster path. 4K mesh video needs the GPU, which is a different piece of work.

Two things that turned out not to be true

Both were expected to matter, and neither did. Recorded because the next person to look at this will have the same two ideas.

Giving each cell only its own slice of the frame made no difference — 11.82 ms against 11.86 ms. The theory was that clipping still hands the whole frame to the transform, so N cells transform it N times. Qt's clip evidently already avoids that work. It also means the Phase 1 image renderer, which draws the whole buffer per cell and clips, is not leaving anything on the table.

Wrapping the frame's own bits without converting is not available. Frames arrive as Format_YUV420P, which no QImage format can wrap. The measurement that suggested it cost 0.05 ms was wrapping YUV bytes in an RGB image and producing garbage quickly. The tool now checks the format and reports "n/a" rather than a number that means nothing.

Recommendation

Build it, for 1080p, using QVideoFrame.paint(), with smoothing exposed per object — but run the tool on the show laptop first. The number that decides it is the one from the machine that has to hold it for two hours, not this one.

What building it actually found

Phase 3 shipped on this basis. Three things the spike did not catch, recorded so the numbers above are read with them in mind.

The frame has to be held, not read back. QVideoSink.videoFrame() is only dependable while videoFrameChanged is being delivered; a moment later the backend may have recycled the buffer, and painting from it draws nothing. MeshVideoLayerItem keeps a reference to the frame it was handed. That is also why it lets go of it on unload — an off-air scene must not pin a full-size buffer.

QVideoFrame.paint() does honour the painter's clip and transform. It was briefly believed not to, on the strength of a render that looked like the warp had come apart. It had not: the test clip was ffmpeg's testsrc2, which contains a magenta bar, and the magenta background being used to find gaps was indistinguishable from it. Twice. Solid-color clips, or the two-background coverage trick in tests/test_mesh_video.py, avoid the trap.

Composing into a buffer first costs about 6.5 ms a frame at 1080p, measured through the real layer rather than the spike's harness. So the shipped renderer draws each cell straight from the frame, and only composes when it must: an aiming grid has to be warped with the picture to be any use, and a vignette applied after the warp would be the wrong shape. Both pay the 6.5 ms, and neither is on during a show.

Through the real layer the per-cell route measures 12.2 ms at one cell, 15.2 at 4x4 and 21.5 at 8x8 — close to the spike. 16x16 comes out at 45.8 ms rather than 27.4, because the spike drew into a bare canvas with no scene graph around it. Keep 1080p mesh video at 8x8 or below on hardware like this.