Distillation for price: USD 0.14/sec at 1080p
xAI did not ship a new architecture here. Instead, the company distilled its flagship Grok Imagine Video 1.5 into a Lite variant, trading some quality for lower cost and faster turnaround. The OpenRouter model card is explicit: Lite trades quality for speed and price, and 1080p output is rendered at 720p and then upscaled.
The pricing is aggressive: USD 0.02/sec at 480p, USD 0.03/sec at 720p, USD 0.14/sec at 1080p. At the same 1080p tier, Veo 3.1 with audio costs USD 0.40/sec, so Lite is roughly 65% cheaper than Veo 3.1 and 44% cheaper than the full Grok Imagine Video 1.5 at USD 0.25/sec. For 60 seconds of total footage, Lite at 1080p is USD 8.40, full 1.5 is USD 15.00, and Veo 3.1 is USD 24.00.
Two ranks above Veo 3.1, on two Pareto frontiers
Artificial Analysis places Lite at No. 17 on both the AA-Video-T2V v2.0 and AA-Video-T2V-Silent v2.0 leaderboards, two places above Veo 3.1. AlphaSignal reporting notes that Lite also lands on the Pareto frontiers for quality-vs-speed and quality-vs-price, meaning that within its band, no measured alternative improves both dimensions at once.
Median generation latency is 60.5 seconds for a 10-second 1080p clip, fast enough to support iterative prompt work and storyboarding without long batch queues. For reference, Kling 3.0 1080p Pro scores slightly higher on quality but takes 94 seconds to produce a 5-second clip; Vidu Q3 Turbo 5-second 720p configuration finishes about nine seconds sooner than Lite 10-second 1080p output but receives a substantially lower quality score.
Strong in multi-scene narrative, weak in lip-sync and anatomy
Across the AA-Video-T2V v2.0 capability groups, Lite comes closest to the frontier in Multi-Scene & Narrative, Lighting & Materials, and Text Rendering. These are also the dimensions where the gap to the full 1.5 model is smallest.
The largest deficits appear in Dialogue & Lip Sync and Human Anatomy. On the use-case axis, Lite performs best on Architecture & Real Estate, Consumer Content, and Productivity & Knowledge Work — exterior fly-throughs, interior walkthroughs, virtual staging, social clips, b-roll, explainers, and educational clips. Live-Action Film and the Frontier category show the largest gaps, so dialogue-heavy scenes and anatomy-sensitive close-ups warrant prompt-specific evaluation before production use.
1080p is upscaled — read the fine print
The OpenRouter model card is explicit that Lite 1080p output is rendered at 720p and upscaled: the frame dimensions are 1080p, but native detail comes from a 720p render. For fine textures, small text, cropping, and compositing, that distinction matters. Some third-party listings expose only 480p and 720p for this model ID, so applications that depend on fine 1080p detail should compare Lite upscaled output against the full model output on representative prompts.
A practical two-tier pipeline
Lite is built for drafts, not replacement. Ten 10-second Lite drafts at 1080p cost USD 14.00; regenerating one selected prompt with the full model adds USD 2.50, for a total of USD 16.50. Ten full-model drafts would cost USD 25.00. Dropping the draft tier to 720p brings the bill to USD 3.00 for ten drafts, plus USD 2.50 for one full-model 1080p generation.
This works because Lite benchmark profile aligns most closely with architectural walkthroughs, explainers, social clips, b-roll, and multi-scene drafts. For synchronized dialogue, anatomy-sensitive close-ups, and demanding live-action sequences, the full 1.5 or a stronger specialist remains the right budget allocation.
Veo 3.1 is the model getting squeezed
The numbers line up: at the same 1080p tier, Lite is roughly 65% cheaper than Veo 3.1 with audio, ranks two places above Veo 3.1 on the AA leaderboard, and renders a 10-second 1080p clip in a 60.5-second median. The mid-band of the video generation market — between basic tools and frontier-priced generators — is exactly where xAI is now applying pressure.
Veo 3.1 has been positioned around audio integration and long-context generation, while Veo 3.1 Fast occupies a cheaper adjacent tier. Lite lands on the Pareto frontier above Veo 3.1 while costing roughly 35% as much. AA-Video-T2V v2.0 rankings depend on a specific prompt set, scoring methodology, and model snapshot, so production conclusions should add route availability, failed-job behavior, queue latency, audio quality, and benchmarks against user own prompts. Lite has done its part: cheapest at its tier, fastest at its tier. The next variable is distribution — whether fal and Vercel AI Gateway can move it from announcement to developer workflows quickly enough.