On July 17 at WAIC, Baidu Yijing debuted a digital-human video podcast solution, pushing generative AI into the "industrial-grade interview content" production scenario. The technical core is three layers of bottom-line moats: Wenxin large model does fine-grained planning of micro-expressions, no longer bound by single emotion labels; the self-developed digital-human specialized model pushes expression control granularity to within 1 second, with eye movement, body, and voice precisely synchronized; massive professional audio-visual sample fine-grained training achieves deep multimodal unity. At the product level, three AI agents — screenwriter, director, and editor — collaborate, paired with 80+ skills to auto-generate speculative interview scripts covering finance, law, technology, and humanities verticals. Officially disclosed measured data: production cost down 74%, content expressiveness up 785%, average play time up 29%. What this solution truly breaks through isn't "human-likeness" but "whether it can consistently do opinion-driven interviews" — a key step for LLM multimodal moving from "performance" to "expression".