[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"news-slug-flexsql-nus-text-to-sql-spider2-65pct-gpt-oss-120b":3,"news-related-0d0e5ce8-fa18-4907-b811-2918ff8464e4":36},{"id":4,"title":5,"summary":6,"content":6,"original_url":7,"source_id":8,"tags":9,"translations":23,"news_slug":29,"published_at":30,"created_at":31,"modified_at":32,"is_published":33,"publish_type":34,"image_url":13,"view_count":35},"0d0e5ce8-fa18-4907-b811-2918ff8464e4","FlexSQL：小型LLM如何在Text-to-SQL任务上超越GPT-o3和DeepSeek-R1","在大模型军备竞赛中，参数量越大能力越强几乎成为共识。但来自新加坡国立大学StringNLPLAB团队的最新研究，正在动摇这一规律。该团队提出FlexSQL，一种Text-to-SQL智能体，其核心设计原则是灵活的数据库交互：智能体可以在推理过程中随时探索模式结构、检查数据值、运行验证查询，而不是像传统系统那样仅在开始时一次性检索模式信息。\n\nFlexSQL生成多样化执行计划以覆盖多种查询解释方式，同时支持SQL和Python两种执行模式，根据任务类型灵活切换。其两层修复机制能够从代码级错误回溯到计划级修订，而传统系统只能在事后修复。\n\n在Spider2-Snow基准测试中，使用gpt-oss-120B的FlexSQL达到了65.4%的得分，超越了使用更强更大模型的GPT-o3和DeepSeek-R1。当FlexSQL作为技能集成到Claude Code中时，实现了超过10%的相对提升。\n\nFlexSQL证明了架构的灵活性可能比模型规模更重要。对于企业部署而言，这意味着可以在保持高性能的同时使用更小、更便宜、更高效的模型，从而显著降低成本。这项工作呼应了近期测试时计算的趋势：给予模型更多的推理时间和交互自由，往往比堆叠参数更有效。","https:\u002F\u002Farxiv.org\u002Fabs\u002F2605.02815","7437aeb9-930c-4866-a2e9-48003c1a792b",[10,14,17,20],{"id":11,"name":12,"slug":12,"description":13,"color":13},"e82b2d09-81b2-43d1-977e-e018443b3c14","coding-agent",null,{"id":15,"name":16,"slug":16,"description":13,"color":13},"0ef8513a-0a26-42f0-b6f9-5b6dadded45c","efficiency",{"id":18,"name":19,"slug":19,"description":13,"color":13},"0a93ec8e-ea39-4693-81de-563ca8c173f7","inference",{"id":21,"name":22,"slug":22,"description":13,"color":13},"01598627-1ea6-4b27-a5d8-874971571a71","llm",[24],{"id":25,"lang":26,"title":27,"summary":28,"content":13},"a3b32365-6f07-4712-a658-0728dcfe0167","en","FlexSQL: small LLMs beat GPT-o3 and DeepSeek-R1 on Text-to-SQL","In the LLM arms race, the conventional wisdom is the bigger the parameters, the stronger the capability. But the latest research from the StringNLPLAB team at the National University of Singapore is shaking this rule. The team proposes FlexSQL, a Text-to-SQL agent whose core design principle is flexible database interaction: the agent can explore schema structures, check data values, and run validation queries at any time during reasoning, rather than retrieving schema information only once at the start like traditional systems.\n\nFlexSQL generates diverse execution plans to cover multiple query interpretations, and supports both SQL and Python execution modes, flexibly switching based on task type. Its two-layer repair mechanism can roll back from code-level errors to plan-level revisions, while traditional systems can only do post-hoc fixes.\n\nIn the Spider2-Snow benchmark, FlexSQL using gpt-oss-120B reached 65.4% score, surpassing GPT-o3 and DeepSeek-R1 which use stronger and larger models. When FlexSQL is integrated as a skill into Claude Code, it achieves over 10% relative improvement.\n\nFlexSQL proves that architectural flexibility may be more important than model scale. For enterprise deployment, this means maintaining high performance while using smaller, cheaper, more efficient models, significantly reducing cost. This work echoes the recent test-time-compute trend: giving the model more reasoning time and interaction freedom often works better than stacking parameters.","flexsql-nus-text-to-sql-spider2-65pct-gpt-oss-120b","2026-05-05T10:15:00Z","2026-05-05T10:12:42.981010Z","2026-08-19T02:08:40.142862Z",true,"agent",104,{"items":37},[38,43,48,53,58,63],{"id":39,"title":40,"news_slug":41,"published_at":42},"4bbc55d2-cabc-477f-a3ad-4e2c119aff2a","TokTier 抓住 Agent 推理的隐藏瓶颈：缓存命中 94.1%，分词仍吃掉 64% 首 token 时间","toktier-stateful-tokenization-agent-serving","2026-07-31T17:56:30+00:00",{"id":44,"title":45,"news_slug":46,"published_at":47},"6b52b4a9-d567-46b8-99c1-e9c65ba59b16","SWE-Pruner Pro:ByteDance 让 Agent 自己当剪枝器,省 39% token 还涨分","swe-pruner-pro-bytedance","2026-07-25T12:00:00+00:00",{"id":49,"title":50,"news_slug":51,"published_at":52},"37aa0bc9-d135-444f-842e-0b40388d29e9","Qwen3.7-Max 原生兼容 Anthropic API 协议：Claude Code 现已可直接调用阿里模型","qwen3-7-max-anthropic-api-claude-code","2026-05-27T10:05:00+00:00",{"id":54,"title":55,"news_slug":56,"published_at":57},"57b691cc-476c-4427-8618-e29127654b34","AMD ROCm 7 原生支持 Qwen3-Coder-Next：单卡 256k 上下文打破推理硬件垄断","amd-rocm7-qwen3-coder-next-256k-mono","2026-05-25T16:10:00+00:00",{"id":59,"title":60,"news_slug":61,"published_at":62},"69e52a42-19c7-4580-8c49-5446233fbdde","7B模型如何超越GPT-4o？ICLR Oral论文揭示AgentFlow流式训练新范式","agentflow-7b-icrl-oral-flow-grpo-14-9pct","2026-05-03T01:10:00+00:00",{"id":64,"title":65,"news_slug":66,"published_at":67},"cac485ab-a429-4ebc-88c6-a1f924f978ff","AWS 开源 KeysAndValues:微调时就让模型学会“遗忘”,单张 A100 撑住 128K","aws-keysvalues-sparse-attention-finetuning","2026-08-26T05:20:00+00:00"]