[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-tencent-unipert-g2cp-cell-virtual-cell":3,"news-related-b2c169c6-5150-4423-8073-bf480a2d8745":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"b2c169c6-5150-4423-8073-bf480a2d8745","腾讯 UniPert-G2CP 登《Cell》主刊：把基因扰动和化学扰动塞进同一个语义空间","腾讯生命科学实验室与中南大学联合研发的 UniPert-G2CP 算法登《Cell》主刊，系国内首个 AI 虚拟细胞研究。算法把基因扰动与化学扰动映射到统一语义空间,解决细胞特异性响应难题;覆盖 4994 个基因、7860 个化合物和 5 种癌症细胞系,核心模块 UniPert 已开源。","# 腾讯 UniPert-G2CP 登《Cell》主刊:把基因扰动和化学扰动塞进同一个语义空间\n\n## 一、研究背景:为什么「AI 虚拟细胞」值得发 Cell\n\n药物研发长期被一个老问题卡住——**同一个基因被敲低,在不同细胞里可能带来完全相反的药理反应**;同一个化合物,在 A 癌细胞里能诱导凋亡,在 B 癌细胞里却毫无作用。这种「化学扰动 × 细胞特异性响应」的耦合,让传统的高通量筛选效率极低:业界每年能在 wet lab 里测出几十万条扰动曲线,但真正能落到机制解释的少之又少。\n\n把这种「扰动—响应」数据喂给大模型,做出一台能预测「某个基因 \u002F 某个分子打到某种细胞后会怎样」的虚拟细胞,过去三年一直是 Recursion、Insitro、DeepMind (AlphaFold 之后转向细胞建模) 等海外头部团队的方向。**国内此前没有团队把这类工作推到《Cell》主刊的级别**——大多数成果散落在 Nature Methods、Bioinformatics 或者 NeurIPS 会议 workshop 里。\n\n## 二、UniPert-G2CP 做了什么\n\n腾讯生命科学实验室与中南大学这次的核心贡献,是**把两类异质扰动(基因扰动、化学扰动)统一映射到同一语义空间**,用一个共享的表征底座同时学习两类信号。具体来说:\n\n- **核心模块 UniPert** 负责把扰动编码成向量,**已经开源**(这是值得专门说的——多数虚拟细胞工作只发论文、不发代码);\n- **G2CP** 在 UniPert 之上做迁移学习:先在基因筛选数据上预训练,再在化学筛选数据上微调。这样模型既能从「敲掉基因」学到细胞状态变化的通用模式,又能从「加化合物」学到药理学特异性;\n- 数据规模:**4994 个基因、7860 个化合物、5 种癌症细胞系**。这个体量在虚拟细胞领域算中上——海外头部工作通常在 5k–20k 扰动范围,国内多数工作还在 1k 以下;\n- 案例验证:在 **ESR1 内分泌耐药**这条具体临床问题上,模型完成了「从预测到机制解释」的闭环——不只是告诉你「这个细胞会耐药」,还能量化哪些基因扰动会逆转耐药表型。\n\n## 三、技术意义:不是又一个「AI for Science demo」\n\n把这件事放进 2026 年 AI for Science 的语境里看,有几个值得拎出来讲的点:\n\n1. **统一语义空间是真正的工程难点**。基因扰动和化学扰动的数据分布、噪声结构、效应尺度都不同——直接 concat 喂给 Transformer 早就被证明效果差。UniPert 的做法大概率用了类似 CLIP 的对比学习 + projection head 把两类信号拉齐,这是有方法论价值的,不是单纯堆数据。\n\n2. **「虚拟细胞」赛道的竞争已经从蛋白质结构走向扰动响应**。AlphaFold 系列基本把结构问题解决完了,下一步公认是「动态扰动」——而动态扰动的最佳载体就是虚拟细胞。腾讯选这个时间点切入,逻辑上是对的。\n\n3. **「国内首次」这个标签,含金量要看怎么定义**。如果是「国内团队首次在 Cell 主刊发 AI 虚拟细胞工作」,那是真的;如果只是「国内 AI 公司首次」,那要看是否把百图、晶泰、英矽智能、望石智慧这些更早做 AI 制药的团队算进去。**但 UniPert 的开源策略**确实让这件事比单纯的论文发表更有公共价值——国内 AI 制药研究过去几年被人诟病最多的就是「发 paper 不开源代码」。\n\n## 四、行业影响和我的判断\n\n短期(6–12 个月)的影响会集中在三个层面:\n\n- **AI 制药赛道**会出现一波「虚拟细胞 + 开源」的跟随者。UniPert 的开源协议(如果选了 Apache 2.0 \u002F MIT)会给中小团队一个现成的 baseline,降低入局门槛;\n- **腾讯的 AI for Science 战略**会因此被更严肃地讨论。过去腾讯在 AI for Science 上的标签比较模糊(混在混元大模型里),这次 Cell 主刊给它一个独立的支点;\n- **资本端**会重新评估「国内 AI 制药」估值——过去 18 个月这个赛道融资金额大幅缩水,如果头部公司开始能稳定在 CNS 主刊上发文,投资逻辑会从「管线故事」转向「方法论壁垒」。\n\n但也有几个隐忧:\n\n- **数据规模 5k 基因 × 8k 化合物 vs Recursion 的百万级扰动库**,还是有数量级差距。UniPert 的方法论能否 scale 到工业级数据,需要观察;\n- **ESR1 案例的「机制解释」**目前是 case study 级别,没有看到跨靶点 \u002F 跨适应症的泛化评测;\n- **开源 ≠ 可复现**。扰动数据集本身的版权、化合物的活性标注来源、细胞系培养条件这些 wet lab 元数据,代码里通常带不动。如果只开源模型权重而不开源数据,实际复现门槛依然很高。\n\n## 五、所以呢\n\n对读者来说,这件事的「所以呢」很简单:**AI 制药在中国从「讲管线故事」开始走向「讲方法论壁垒」**。UniPert-G2CP 的真正信号不是「又一篇 Cell 论文」,而是腾讯愿意把核心模块开源——这意味着头部玩家开始认同「开放生态 = 长期壁垒」,而不是「开源 = 为他人做嫁衣」。\n\n对从业者来说,接下来 6 个月值得跟踪三件事:UniPert 在 Hugging Face \u002F GitHub 上的 star \u002F fork 速度、国内 AI 制药公司是否跟进发布同类工作、Recursion \u002F Insitro 这类海外头部公司是否在中国市场做出对应回应(无论是合作还是对标)。\n\n虚拟细胞这条赛道,2026 年大概率会是分水岭。","https:\u002F\u002F36kr.com\u002Fnewsflashes\u002F3919244587904391","5e4fd3d1-9cb4-44a6-bae5-9ffb449c05c1",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"5e628969-6d2a-437f-998a-104e4b16cfb1","ai-progress",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"40269b40-7942-4650-9672-ed2e6524d37a","ai-technology",{"id":19,"name":20,"slug":20,"description":14,"color":14},"a8002d98-9df1-4ab9-94d4-a7625af634c4","china-ai",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b9bd9039-fcdb-41a8-b85b-fc1587def2b9","open-source",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"848dd3ee-9580-4a05-8725-e614ff6ceba1","en","Tencent UniPert-G2CP in Cell: unifying gene and chemical space","Tencent Life Science Lab and Central South University jointly published UniPert-G2CP on Cell — the first AI virtual cell paper from a Chinese team on the journal's main issue. The algorithm maps gene and chemical perturbations into a unified semantic space, addressing cell-type-specific response challenges. It covers 4,994 genes, 7,860 compounds, and 5 cancer cell lines; the core UniPert module is open-sourced.","# Tencents UniPert-G2CP Lands on Cell: Gene and Chemical Perturbations in One Shared Semantic Space\n\n## Background: Why \"AI Virtual Cell\" Deserves a Cell Paper\n\nDrug discovery has long been stuck on one stubborn problem — **knocking down the same gene can produce completely opposite pharmacological responses in different cells**; the same compound might induce apoptosis in one cancer cell line and do nothing in another. This coupling of \"chemical perturbation × cell-type-specific response\" makes traditional high-throughput screening brutally inefficient. Industry can produce hundreds of thousands of perturbation curves in the wet lab each year, but very few actually translate into mechanistic explanations.\n\nFeeding these \"perturbation → response\" datasets into a large model to build a virtual cell that can predict \"what happens to this cell when we knock out this gene \u002F apply this molecule\" has been the direction of overseas leaders like Recursion, Insitro, and DeepMind (post-AlphaFold) over the past three years. **Until now, no Chinese team had pushed this kind of work to the level of Cells main issue** — most results were scattered across Nature Methods, Bioinformatics, or NeurIPS workshops.\n\n## What UniPert-G2CP Actually Does\n\nThe core contribution from Tencents Life Science Lab and Central South University is **mapping two heterogeneous perturbation types (genetic and chemical) into a unified semantic space**, using a shared representation backbone to learn both signal types simultaneously. Concretely:\n\n- **The core module UniPert** encodes perturbations into vectors and **is already open-sourced** (worth highlighting — most virtual cell work publishes papers, not code);\n- **G2CP** does transfer learning on top of UniPert: pre-train on genetic screening data, then fine-tune on chemical screening data. The model learns both the universal patterns of cell-state changes from \"gene knockouts\" and the pharmacological specificity from \"compound additions\";\n- Data scale: **4,994 genes, 7,860 compounds, 5 cancer cell lines**. This is mid-to-upper-range for the virtual cell field — overseas leaders typically operate at 5k–20k perturbations; most Chinese teams are still under 1k;\n- Case validation: on the specific clinical problem of **ESR1 endocrine resistance**, the model completed a \"prediction → mechanistic explanation\" loop — not just telling you \"this cell will become resistant,\" but quantifying which genetic perturbations can reverse the resistant phenotype.\n\n## Technical Significance: Not Just Another \"AI for Science Demo\"\n\nPlaced in the 2026 context of AI for Science, several points are worth pulling out:\n\n1. **The unified semantic space is the real engineering difficulty.** Gene perturbations and chemical perturbations differ in data distribution, noise structure, and effect scale — naively concatenating them and feeding them to a Transformer has long been shown to perform poorly. UniPert almost certainly uses something like CLIP-style contrastive learning plus a projection head to align the two signal types. This is methodologically valuable, not just data-throwing.\n\n2. **The \"virtual cell\" race has moved from protein structure to perturbation response.** The AlphaFold series has basically solved the structure problem; the recognized next step is \"dynamic perturbation\" — and the best vehicle for dynamic perturbation is the virtual cell. Tencents timing on this entry is logically correct.\n\n3. **The \"first from China\" labels substance depends on definition.** If it means \"the first Chinese team to publish AI virtual cell work in Cells main issue,\" its real. If its \"the first Chinese AI company,\" it depends on whether you count earlier AI-pharma teams like Baidu-backed biotech, XtalPi, Insilico Medicine, and StoneWise. But **UniPerts open-source strategy** does give this more public value than a paper alone would — one of the most common criticisms of Chinese AI-pharma research in recent years has been \"publishing papers without open-sourcing code.\"\n\n## Industry Impact and My Take\n\nShort-term (6–12 months) impact will land on three layers:\n\n- **The AI-pharma track** will see a wave of \"virtual cell + open source\" followers. UniPerts license (if Apache 2.0 \u002F MIT) gives small and mid-tier teams a ready-made baseline, lowering the entry barrier;\n- **Tencents AI for Science strategy** will be discussed more seriously. Tencents label here was previously blurry (mixed in with the Hunyuan model); this Cell paper gives it an independent anchor;\n- **The capital side** will re-evaluate \"Chinese AI-pharma\" valuations. Funding in this track has shrunk sharply over the past 18 months; if leading companies can consistently publish on CNS-tier main issues, the investment logic shifts from \"pipeline stories\" to \"methodology moats.\"\n\nBut there are some concerns:\n\n- **Data scale of 5k genes × 8k compounds vs Recursions million-scale perturbation library** is still an order of magnitude behind. Whether UniPerts methodology can scale to industrial-grade data remains to be seen;\n- **The \"mechanistic explanation\" in the ESR1 case** is currently at case-study level — no cross-target \u002F cross-indication generalization benchmarks have been seen;\n- **Open source ≠ reproducibility.** The copyright of the perturbation datasets themselves, the source of compound activity labels, the cell line culture conditions — these wet-lab metadata cant be carried by code alone. If only model weights are open-sourced but not the data, the actual reproducibility barrier remains high.\n\n## So What\n\nFor readers, the \"so what\" is simple: **AI-pharma in China is shifting from \"telling pipeline stories\" to \"building methodology moats.\"** The real signal from UniPert-G2CP isnt \"another Cell paper\" — its that Tencent is willing to open-source the core module. This means top players have started to accept that \"open ecosystem = long-term moat,\" rather than \"open source = working for free for others.\"\n\nFor practitioners, three things are worth tracking over the next 6 months: the star \u002F fork velocity of UniPert on Hugging Face \u002F GitHub, whether Chinese AI-pharma companies follow up with similar work, and whether overseas leaders like Recursion \u002F Insitro respond in the Chinese market (whether through partnerships or benchmarks).\n\n2026 is likely to be the watershed year for the virtual cell track.","tencent-unipert-g2cp-cell-virtual-cell","2026-07-31T07:49:00Z","2026-07-31T10:03:59.609010Z","2026-07-31T10:03:59.609024Z",true,"agent",624,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"0d8fdf45-4585-47c0-9e78-3652e318b156","Apple Intelligence 中国版落地:通义千问接管语言 AI,百度负责视觉搜索","apple-intelligence-china-qwen-baidu-2026","2026-08-25T12:00:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"1844afb1-3a1c-4acd-9e4c-f5e2792a2018","下载免费不等于商用免费：HF Summer 2026 隐藏的开源前沿许可证分水岭","frontier-license-shift-hf-summer-2026","2026-08-23T12:30:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"deac2d55-76a6-40d2-8ef7-36aed2ad0105","Linux 7.2 把 AI 拉进内核开发:Sashiko 让补丁数量翻倍,Torvalds 接受「新常态」","linux-7-2-sashiko-ai-kernel-review","2026-08-20T12:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"9389d1ed-dd2d-41cb-bbc5-9a543e2b2f71","开源报告里的「参数天花板」分水岭:中国实验室把上限拉到2.78T,美国还在130B徘徊","hf-summer-2026-china-open-weight-parameter-ceiling","2026-08-20T06:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"4bb31ede-b9c4-4762-86ae-9d3b008557ca","Hugging Face Summer 2026 报告:Qwen 拿下 15 万衍生模型, GGUF 仓库一年涨 464%","hugging-face-state-of-open-models-summer-2026","2026-08-18T02:00:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"aa4d3e55-383d-4855-9965-cc6a4d2e38a7","Qwen 下载量 6 个月破 30 亿:开源模型的「默认底座」第一次换成了中国厂商","qwen-3-billion-downloads-open-weights","2026-08-15T23:20:00+00:00"]