On July 7, Ant Group's Lingbo Tech released the spatial perception model LingBot-Depth 2.0 along with the companion vision base LingBot-Vision. Based on 150 million-scale training data, this generation pushes the boundary of the "robot's eye" another circle out. The official points out four upgrade directions: edge clarity, small-object recognition, long-distance depth estimation, and complex-scene robustness. Depth estimation looks like a classic CV task, but when it lands on embodied robots, any blur point can make a grasp fail. After LingBot-Vision's simultaneous release, LingBot upgrades from a single depth model to a "base + task" dual-layer architecture, echoing the industry paradigm of "general vision base + downstream task model". Zooming out a bit: from UFP4 quantization, the Ring-2.6-1T trillion-thinking model, to this LingBot-Depth 2.0, Ant's Bailing/Lingbo two lines have formed a clear "base + embodied" dual-line strategy. LLMs fight for the general-intelligence ceiling, embodied vision fights for the physical-world entry — two curves are being pulled up simultaneously by Chinese big tech. LingBot-Depth 2.0 didn't publish a paper or benchmark numbers, but Ant taking "see it → see it accurately" as the external positioning means they've treated precision as the core metric for the second-half competition, rather than yet another leaderboard-climbing PR battle.