[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-base44-base-1-vibe-coding-llm-launch":3,"news-related-dea38861-2618-4468-9bad-a18eea96a818":38},{"id":4,"title":5,"summary":6,"content":7,"original_url":8,"source_id":9,"tags":10,"translations":24,"news_slug":31,"published_at":32,"created_at":33,"modified_at":34,"is_published":35,"publish_type":36,"image_url":14,"view_count":37},"dea38861-2618-4468-9bad-a18eea96a818","Base44 Base 1：年入 1.5 亿美元的 vibe-coding 平台，终于把自己的 LLM 训出来","Base44(Wix 子公司,ARR 1.5 亿美元)推出自研 LLM Base 1,基于数千万次真实应用构建交互训练,是首个 vibe-coding 平台内部自研模型,直接对外可用并自动路由生产流量。","## 从调用别人模型,到自己拥有模型\n\nBase44——Wix 旗下、ARR 1.5 亿美元的 vibe-coding 平台——在 7 月初悄悄交出了自己训练的第一款 LLM:**Base 1**。这件事本身在开源\u002F前沿模型满天飞的 7 月并不起眼,但它指向了一个更值得讨论的趋势:当一个垂直场景足够「富」,平台方自己训一个模型的拐点已经到了。\n\n## Base 1 是什么\n\n- **训练数据**:数千万次真实的应用构建交互(用户描述需求→模型写代码→用户修改→再次生成),全部来自 Base44 平台内部。\n- **定位**:**Vibe-coding 专用的 LLM**,不是通用对话模型。它不试图跟 GPT-5.5、Claude Opus 4.8 比「世界知识」,它只比一件事——把一句自然语言变成一个能跑的全栈应用。\n- **部署方式**:内部 auto-routing 已经会根据内部 benchmark 胜出率把任务直接路由到 Base 1,而非外部 API。\n- **战略意义**:Base44 是第一个把「自己训的 vibe-coding LLM」作为产品卖点公开喊出来的平台(连 Bolt、Lovable、Replit Agent 都还停留在「接最强闭源 API」)。\n\n## 为什么这件事值得拆\n\n**1. 数据飞轮。** Base44 的核心资产不是模型,是「用户在做 app 时怎么改、怎么回滚、卡在哪、最后交付什么」这数千万条交互轨迹。这些数据**任何外部模型都拿不到**,而且每过一天都在增加。这条护城河是 0 边际成本的。\n\n**2. 闭源模型厂商的「看不见的天花板」。** 当场景足够垂直,通用模型的能力溢出越来越严重——参数规模、上下文长度、世界知识,对一个只想生成 CRUD 表单+鉴权+部署的小应用来说全是浪费。Base 1 的目标函数比 GPT-5.5 简单得多:在 vibe-coding 这一项上超过所有外部模型,其他全部不要。\n\n**3. 单位经济学。** Base44 ARR 1.5 亿美元,如果 50% 的 token 成本从 Claude\u002FGPT 切到 Base 1(自托管),按外部 API 大约 3\u002F15 美元\u002F百万 token 的混合价、自托管 GPU 0.6 美元\u002F百万 token 估算,光是这一项一年就能省下 2000–3000 万美元。这还不算延迟下降带来的转化率提升。\n\n**4. 验证了「小而专」的工程哲学。** 与其卷一个 1T+ 的通用模型去追 GPT,不如在一个 50B 量级、把整个 1.5 亿 ARR 业务闭环的窄域任务上做到 95 分。这跟 Cognition 的 SWE-1.7(基于 Kimi K2.7 后训练、专攻 Devin 任务)是同一种思路——只不过 SWE-1.7 走开源后训练,Base 1 走全自研。\n\n## 这对其他 vibe-coding 平台意味着什么\n\n- **Bolt \u002F Lovable \u002F v0**:如果不在 6–12 个月内拿出自己的模型(或找到独家数据合作),单位经济会被 Base44 拉开 1–2 个量级。\n- **Replit Agent \u002F Cursor Composer**:Replit 已经在押 CodeSandbox + 自家 agent 框架,Cursor 走 Composer 自研路线;Base44 的动作会加速这两家把「自研」从 P0 项目提升到 P0+。\n- **闭源 API 厂商**:这是 Anthropic \u002F OpenAI 第一个「看不顺眼但也挡不住」的客户流失形态——平台方不是不想用,是要把自己的利润留在自己手里。\n\n## 我的判断\n\nBase 1 不会是 7 月最强的 LLM,但它可能是 7 月**最容易被低估**的:一个 ARR 1.5 亿美元的垂直平台,正式宣告「我的 LLM 比你的 LLM 更懂我的用户」。这种「场景富→数据飞轮→自研模型→单位经济反超」的闭环,正是过去两年 SaaS 行业反复念叨却很少有公司真正跑通的剧本。Base44 跑通了。\n\n下一个值得盯的信号:Base44 会不会在 3 个月内,把 Base 1 通过 API 卖给第三方 vibe-coding 工具(类似 Cognition 把 SWE-1.7 留在内部,但不排除 API)。如果走通,vibe-coding 这条赛道会出现第一个「既卖铲子又挖金子」的平台。","https:\u002F\u002Fthursdai.news\u002Freleases\u002F2026-07","3bd971a8-3897-43d9-84ac-43879efd2f94",[11,15,18,21],{"id":12,"name":13,"slug":13,"description":14,"color":14},"e82b2d09-81b2-43d1-977e-e018443b3c14","coding-agent",null,{"id":16,"name":17,"slug":17,"description":14,"color":14},"01598627-1ea6-4b27-a5d8-874971571a71","llm",{"id":19,"name":20,"slug":20,"description":14,"color":14},"7e89b5cc-57db-4f37-bc6d-28919a73931c","model-release",{"id":22,"name":23,"slug":23,"description":14,"color":14},"b1853a5a-d940-42b7-94f9-0488ee3f2cf7","new-model",[25],{"id":26,"lang":27,"title":28,"summary":29,"content":30},"6f54acd1-958f-4779-9ef4-3aaf63d62071","en","Base 44 trains its own LLM behind a $150M vibe-coding run","Base44, a Wix subsidiary with $150M ARR, launched its in-house LLM Base 1 — trained on tens of millions of real app-building interactions, the first vibe-coding platform to ship a proprietary model, with auto-routing already directing production traffic to it.","## From calling someone else's model, to owning your own\n\nBase44 — the Wix subsidiary running a $150M-ARR vibe-coding platform — quietly shipped its first self-trained LLM in early July: **Base 1**. The release barely registered against the front-page model launches of the month, but it points to a trend worth taking seriously: when a vertical scenario is rich enough, the inflection point for a platform to train its own model has arrived.\n\n## What Base 1 is\n\n- **Training data:** tens of millions of real app-building interactions (user describes a need → model writes code → user edits → regenerates), all sourced from inside Base44's own platform.\n- **Positioning:** a **vibe-coding-specific LLM**, not a general-purpose chat model. It doesn't try to beat GPT-5.5 or Claude Opus 4.8 on world knowledge — it competes on exactly one thing: turning a natural-language prompt into a running full-stack application.\n- **Deployment:** internal auto-routing already directs tasks to Base 1 whenever it beats alternatives on internal benchmarks, rather than calling an external API.\n- **Strategic significance:** Base44 is the first vibe-coding platform to publicly frame \"our own trained LLM\" as a product differentiator — Bolt, Lovable, and Replit Agent are still parked on \"wire up the strongest closed-source API.\"\n\n## Why this is worth unpacking\n\n**1. The data flywheel.** Base44's core asset isn't the model — it's the tens of millions of trajectories of *how users build, where they get stuck, what they revert, what they ship*. No external model can access this, and it grows for free every day. The moat compounds at zero marginal cost.\n\n**2. The invisible ceiling closed-model vendors can't see.** When a scenario is vertical enough, the capability overflow of general models gets worse over time — parameter count, context length, world knowledge are all waste for a system whose only job is to generate a CRUD form plus auth plus a deploy. Base 1's loss function is dramatically simpler than GPT-5.5's: beat every external model on vibe-coding, ignore everything else.\n\n**3. Unit economics.** Base44's $150M ARR is the kind of business where 50% of token cost migrating from Claude\u002FGPT to Base 1 (self-hosted) — at a rough external-API blended price of $3\u002F$15 per million tokens vs. ~$0.6 per million for self-hosted GPU — saves $20–30M a year on this line item alone. Latency improvements and the resulting conversion lift are on top of that.\n\n**4. A validation of \"small and focused\" engineering philosophy.** Instead of chasing GPT with a 1T+ general model, Base 1 bets on a ~50B-class model that closes the loop on a $150M business in a narrow task. Same play as Cognition's SWE-1.7 (post-trained on Kimi K2.7, scoped to Devin tasks) — except SWE-1.7 went open-weight post-training while Base 1 went fully proprietary.\n\n## What this means for the other vibe-coding players\n\n- **Bolt \u002F Lovable \u002F v0:** if they don't ship their own model (or lock down an exclusive data partnership) within 6–12 months, Base44 will pull 1–2 orders of magnitude ahead on unit economics.\n- **Replit Agent \u002F Cursor Composer:** Replit is already betting on CodeSandbox + its own agent framework; Cursor is going down the Composer path. Base44's move will push both of these to elevate \"self-trained model\" from a P0 project to P0+.\n- **Closed-source API vendors:** This is the first form of customer churn Anthropic \u002F OpenAI can't quite push back on — these platforms don't *want* to leave, they want to keep the margin.\n\n## My read\n\nBase 1 won't be July's strongest LLM, but it may be July's **most underestimated** one: a $150M-ARR vertical platform formally declaring \"my LLM understands my users better than yours does.\" The closed loop — scenario richness → data flywheel → in-house model → unit-economics reversal — is a script the SaaS industry has been reciting for two years but rarely executes. Base44 just did.\n\nThe next signal to watch: will Base44 sell Base 1 through an API to third-party vibe-coding tools within three months (similar to how Cognition kept SWE-1.7 in-house, but APIs are not out of the question)? If it does, vibe-coding gets its first platform that \"both sells the shovels and digs the gold.\"\n","base44-base-1-vibe-coding-llm-launch","2026-07-29T06:00:00Z","2026-07-29T14:11:13.576827Z","2026-07-29T14:11:13.576835Z",true,"agent",58,{"items":39},[40,45,50,55,60,65],{"id":41,"title":42,"news_slug":43,"published_at":44},"89a79f9a-bfd2-4ebe-8f03-92fa74a3a34f","Ornith-1.5 开源：模型自己出题、自己搭考场，397B 到 9B 三档齐发","ornith-1-5-self-improvement-open-models","2026-08-20T13:30:00+00:00",{"id":46,"title":47,"news_slug":48,"published_at":49},"f8a3ead9-b1aa-4f6a-9146-19c01f7f1375","微软把编码模型价格砍到四分之一：138B\u002F5B 稀疏 MoE 加原生视觉全量进入 Copilot","mai-code-1-1-flash-copilot-moe-vision","2026-08-12T06:00:00+00:00",{"id":51,"title":52,"news_slug":53,"published_at":54},"e846f1c6-4644-4f84-a664-83ec82734210","Meta Muse Spark 1.2 与 Muse Code 把「1.2 → 编程」的推理效率推回前沿","meta-muse-spark-12-coding-agent-54-index","2026-08-05T08:00:00+00:00",{"id":56,"title":57,"news_slug":58,"published_at":59},"70b5b0d6-ce28-48e8-abe4-6a667a723c4e","xAI 把 Colossus 推到 2 GW:555,000 颗 GPU 撑起 Grok 4.6\u002F4.7 的万亿参数竞速","xai-colossus-2gw-grok-4-6-7-compute","2026-07-31T04:00:00+00:00",{"id":61,"title":62,"news_slug":63,"published_at":64},"f4af1e4e-98a3-4810-831c-699ffe31ae73","马斯克公布 Grok 4.6\u002F4.7 路线图：1.5T\u002F2.1T 参数，SFT+RL 升级，8 月 7 日发行","grok-4-6-4-7-roadmap-1-5t-2-1t","2026-07-30T08:45:00+00:00",{"id":66,"title":67,"news_slug":68,"published_at":69},"1e7d0673-aecc-42b5-8560-92a2b4d4daf6","快手 KAT-Coder-V2.5 把 Agentic Coding 训练改写成基础设施工程","kuaishou-kat-coder-v2-5","2026-07-27T06:00:00+00:00"]