{"allowContribute":false,"item":{"activeCommit":"5d845135e14f041fa37b734da6b486b6156f8258","apis":{"routes":[],"types":{}},"backendSchedules":null,"bundleHash":"0620fcc394984467","createdAt":1786329812722,"displayName":"Hermes 深度搜尋","gitCommit":"5d845135e14f041fa37b734da6b486b6156f8258","id":"530beed6321917638128ccbc","isPublic":true,"itemType":"PLUGIN","name":"Hermes 深度搜尋","pluginDir":"530beed6321917638128ccbc","repositoryID":"repo_34868767f907c5eae0bd52f1","status":"active","updatedAt":1786332554999,"updatedBy":{"userId":"6a3ddc490146560bbf360f","userName":"Raphael Chan"},"version":9},"ownerMemberId":"6a3ddc490146560bbf360f","ownerName":"Raphael Chan","publicPlugins":{"PLUGIN":{"pluginItemId":"530beed6321917638128ccbc","pluginDir":"530beed6321917638128ccbc","bundleHash":"0620fcc394984467","live":true}},"subtree":[{"activeCommit":"5d845135e14f041fa37b734da6b486b6156f8258","apis":{"routes":[],"types":{}},"backendSchedules":null,"bundleHash":"0620fcc394984467","createdAt":1786329812722,"displayName":"Hermes 深度搜尋","gitCommit":"5d845135e14f041fa37b734da6b486b6156f8258","id":"530beed6321917638128ccbc","isPublic":true,"itemType":"PLUGIN","name":"Hermes 深度搜尋","pluginDir":"530beed6321917638128ccbc","repositoryID":"repo_34868767f907c5eae0bd52f1","status":"active","updatedAt":1786332554999,"updatedBy":{"userId":"6a3ddc490146560bbf360f","userName":"Raphael Chan"},"version":9},{"createdAt":1786329814421,"deletedAt":null,"id":"default_deep_research_folder_6a3ddc490146560bbf360f","isNew":false,"isPublic":false,"itemType":"DEEP_RESEARCH_FOLDER","name":"我的研究","originPluginDir":"530beed6321917638128ccbc","originPluginID":"530beed6321917638128ccbc","parents":{"deepResearch":1786329814421},"preParentID":null,"reviewedAt":0,"updatedAt":1786329814421,"updatedBy":{"userId":"6a3ddc490146560bbf360f","userName":"Raphael Chan"},"version":1},{"confidence":"high","createdAt":1786332167351,"deletedAt":null,"depth":"deep","id":"6a794407fd5f30e90ec156ec","isPublic":false,"itemType":"DEEP_RESEARCH_REPORT","name":"deepseek同kimi 既智力比較","originPluginDir":"530beed6321917638128ccbc","originPluginID":"530beed6321917638128ccbc","outputFormat":"report","parents":{"default_deep_research_folder_6a3ddc490146560bbf360f":1786332167351},"preParentID":null,"query":"deepseek同kimi 既智力比較","reportMarkdown":"## 深度研究：DeepSeek vs Kimi 智力比較（2026 年 8 月）\n\n### 任務分解\n1. **DeepSeek 最新旗艦模型**（V4 Pro）的能力與基準成績\n2. **Kimi 最新旗艦模型**（K3）的能力與基準成績\n3. **第三方頭對頭比較**：智力指數、編程、推理、知識的差異\n4. **獨立評測與反證**：自報 vs 獨立測量、限制與爭議\n\n研究時點：2026-08-10，聚焦兩家目前最新旗艦（DeepSeek V4 Pro、Kimi K3）。\n\n### ✅ 已確認事實\n\n**1. 兩款旗艦基本規格**\n- **DeepSeek V4 Pro**：2026-04-24 發布（7-20 轉正式版），1.6T 總參數 / 49B 活躍參數（MoE），MIT 開源權重，1M token 上下文，純文字（無多模態）— [chinaaibench.com](https://chinaaibench.com/blog/deepseek-v4-review-2026/)\n- **Kimi K3**：2026-07 發布（約 7/16），2.8T 總參數 / 約 50B 活躍參數（MoE），原生多模態，1M 上下文，是目前**最大的開源權重模型**；權重採 modified MIT，原訂 7-27 開放下載 — [deepinfra.com](https://deepinfra.com/blog/kimi-k3-vs-deepseek-v4-pro-vs-glm-5-2)、[cnbc.com](https://www.cnbc.com/2026/07/17/moonshot-ai-kimi-k3-model-openai-anthropic-china.html)\n\n**2. Artificial Analysis 智力指數（第三方統整，最可信的橫向指標）**\n- **Kimi K3：57 分，全球 #3~#4**（2026-08-07 排行榜 #4、57.1%；緊追 Claude Fable 5 的 59.9 與 GPT-5.6 Sol 的 58.9）— [benchlm.ai 排行榜](https://benchlm.ai/benchmarks/artificialanalysis)、[artificialanalysis.ai](https://artificialanalysis.ai/models/comparisons/kimi-k3-vs-deepseek-v4-pro)\n- **DeepSeek V4 Pro：44~45 分（Max reasoning），全球約 #21**；V4 Flash（0731 版）49.9 分、#17 — 同一智力指數 [artificialanalysis.ai](https://artificialanalysis.ai/models/comparisons/kimi-k3-vs-deepseek-v4-pro)、[benchlm.ai](https://benchlm.ai/benchmarks/artificialanalysis)\n- **兩者差距約 13 分**，Kimi K3 明顯領先；K3 也被定位為「智力接近 Opus 4.8 / GPT-5.5，但仍落後 Fable 5 與 GPT-5.6 Sol」— [linkedin.com](https://www.linkedin.com/pulse/kimi-k3-launches-57-artificial-analysis-intelligence-clwlc)、[cnbc.com](https://www.cnbc.com/2026/07/17/moonshot-ai-kimi-k3-model-openai-anthropic-china.html)\n\n**3. 編程能力**\n- **Kimi K3 拿下 LMArena Frontend Code Arena 全球第一**（Elo 1679，勝過 Claude Fable 5 的 1631 與 GPT-5.6 Sol 的 1618），是**首個登上主要編程排行榜榜首的中國模型** — [pasqualepillitteri.it](https://pasqualepillitteri.it/en/news/8475/kimi-k3-jailbreak-pliny-fact-check)、[tomshardware.com](https://www.tomshardware.com/tech-industry/artificial-intelligence/moonshot-releases-2-8-trillion-parameter-kimi-k3)\n- **DeepSeek V4 Pro 在 SWE-bench Verified 較高（80.6% vs K3 的 76.8%）**、LiveCodeBench 93.5% 居全球第一 — [deepinfra.com](https://deepinfra.com/blog/kimi-k3-vs-deepseek-v4-pro-vs-glm-5-2)、[chinaaibench.com](https://chinaaibench.com/blog/deepseek-v4-review-2026/)\n\n**4. Agentic 能力**\n- BenchLM 共享基準全部由 **Kimi K3 勝出**：Terminal-Bench 2.0（K3 88.3% vs V4 Pro 59.1%）、MCP Atlas（84.2% vs 69.4%）、GPQA（93.5% vs 72.9%）、HLE（56% vs 7.7%）— [benchlm.ai](https://benchlm.ai/compare/deepseek-v4-pro-vs-kimi-3)\n- DeepInfra 同樣指出 K3 agentic 評分 89.5 vs V4 Pro 59.1 — [deepinfra.com](https://deepinfra.com/blog/kimi-k3-vs-deepseek-v4-pro-vs-glm-5-2)\n\n**5. 獨立機構評測（NIST CAISI）**\n- CAISI 獨立重測後判斷 DeepSeek V4 Pro **約落後美國前沿 8 個月**（約 GPT-5 水準）；強項是數學、軟體工程、自然科學，弱項是**新情境推理**（ARC-AGI-2、PortBench、CTF）— [geotoolbox.ai](https://geotoolbox.ai/blog/deepseek-v4)、[chinaaibench.com](https://chinaaibench.com/blog/deepseek-v4-review-2026/)\n- 英國 AISI + 美國 CAISI 聯合評測給 **Kimi K3 最高的網絡安全能力分數** — [linkedin.com](https://www.linkedin.com/posts/marialuisaredondo_uk-aisi-caisi-preliminary-assessment-of-activity-7486167583591751680-of0A)\n\n### 🔶 推論／聚合\n- **整體智力結論：Kimi K3 顯著優於 DeepSeek V4 Pro。** 多個獨立來源一致：Artificial Analysis 智力指數 K3 領先約 13 分（57 vs 44）；BenchLM 五項共享基準全數 K3 勝出；agentic 能力差距尤其大（Terminal-Bench 89.5 vs 59.1）。\n- **但這不是一面倒**：DeepSeek V4 Pro 在 SWE-bench Verified（80.6 vs 76.8）與 LiveCodeBench（93.5%）反而領先 K3，且**成本遠低**（V4 Pro $0.435/$0.87 每百萬 token、V4 Flash 約 3 美分/任務；K3 約 $2.31/百萬 token）。挑「純編程任務」時 DeepSeek 是極高 CP 值選擇；挑「綜合智力 / 智能體 / 前端」時 K3 勝。\n- **定位差別**：K3 目標是「open frontier intelligence」，與美國頂級閉源模型同級競爭（僅略遜 Fable 5 / GPT-5.6 Sol）；DeepSeek V4 定位是「追趕型 + 成本殺手」，靠 75% 降價與低價 Flash 搶市占。\n\n### ❓ 待驗證\n- **廠商自報 vs 獨立測量的落差**：DeepSeek 官方宣稱 GPQA Diamond 90.1%、HLE 37.7%，但 BenchLM 獨立量測僅 72.9% / 7.7%，落差極大；NIST CAISI 也直言 DeepSeek 自評高於實測。K3 的 93.5% GPQA 也以廠商自報為主，尚缺等量級的獨立複測。\n- **「K3 擊敗 Fable 5」是否過譽**：該勝績僅限 LMArena 前端代碼垂直領域，LMArena 本質是「人類偏好」評分（漂亮不一定正確）；在 GDPval v2 長程 agentic 任務上 K3（1668）仍落後 Fable 5（1760）約 90 分 — [pasqualepillitteri.it](https://pasqualepillitteri.it/en/news/8475/kimi-k3-jailbreak-pliny-fact-check)\n- **K3 權重與安全資料**：發布時權重尚未立刻可下載（原訂 7-27）、官方未附安全章節 / red-team 結果；Pliny「解放」K3 的越獄宣稱也無第三方複現 — [techi.com](https://www.techi.com/kimi-k3-scale-efficiency-deployment-boundaries/)、[pasqualepillitteri.it](https://pasqualepillitteri.it/en/news/8475/kimi-k3-jailbreak-pliny-fact-check)\n\n### 信心度評估\n- **整體信心度：高**\n- 強力支持：K3 綜合智力領先 V4 Pro 這個結論，由 Artificial Analysis（權威第三方統整）、BenchLM、DeepInfra、CNBC、NIST CAISI 等多個獨立來源交叉一致；差距幅度（約 13 分智力指數）也互相吻合。\n- 主要缺口：**特定基準的絕對數字**（尤其 DeepSeek 自報 vs 獨立測量的差異）可信度偏低，編程 / agentic 細項成績隨測量方法變動；K3 的多項數字仍屬廠商自報待獨立複測。這不影響「誰比較聰明」的整體排序，但影響「差多少」。\n\n### 來源\n1. [Kimi K3 Tech Blog: Open Frontier Intelligence](https://www.kimi.com/blog/kimi-k3) — kimi.com\n2. [DeepSeek V4 review 2026: benchmarks, pricing](https://chinaaibench.com/blog/deepseek-v4-review-2026/) — chinaaibench.com\n3. [DeepSeek V4 Pro vs Kimi K3: Benchmarks \u0026 Cost](https://benchlm.ai/compare/deepseek-v4-pro-vs-kimi-3) — benchlm.ai\n4. [Artificial Analysis Intelligence Index Leaderboard](https://benchlm.ai/benchmarks/artificialanalysis) — benchlm.ai\n5. [Kimi K3 (max) vs DeepSeek V4 Pro (Reasoning, Max Effort)](https://artificialanalysis.ai/models/comparisons/kimi-k3-vs-deepseek-v4-pro) — artificialanalysis.ai\n6. [Kimi K3 vs DeepSeek V4 Pro vs GLM-5.2](https://deepinfra.com/blog/kimi-k3-vs-deepseek-v4-pro-vs-glm-5-2) — deepinfra.com\n7. [China's Moonshot AI unveils Kimi K3](https://www.cnbc.com/2026/07/17/moonshot-ai-kimi-k3-model-openai-anthropic-china.html) — cnbc.com\n8. [DeepSeek V4: What It Actually Is Now](https://geotoolbox.ai/blog/deepseek-v4) — geotoolbox.ai\n9. [Kimi K3 'Liberated' by Pliny: What the Benchmarks Actually Say](https://pasqualepillitteri.it/en/news/8475/kimi-k3-jailbreak-pliny-fact-check) — pasqualepillitteri.it\n10. [Kimi K3 achieves #3 in the Artificial Analysis Intelligence Index](https://www.linkedin.com/pulse/kimi-k3-launches-57-artificial-analysis-intelligence-clwlc) — linkedin.com\n11. [China's 2.8T-parameter Kimi K3 beats Claude Fable 5 in Frontend Code Arena](https://www.tomshardware.com/tech-industry/artificial-intelligence/moonshot-releases-2-8-trillion-parameter-kimi-k3) — tomshardware.com\n12. [UK AISI/CAISI joint assessment of Kimi K3](https://www.linkedin.com/posts/marialuisaredondo_uk-aisi-caisi-preliminary-assessment-of-activity-7486167583591751680-of0A) — linkedin.com\n\n*研究時點 2026-08-10；重點追蹤 DeepSeek V4 系列（2026-04 起）與 Kimi K3（2026-07 起）。*","sourcesCount":12,"starred":false,"status":"completed","taskCategory":"broad_topic_research","updatedAt":1786332458783,"updatedBy":{"agentId":"c71b4a7bc16f7cd8baae9a13","agentName":"Hermes 深度研究員","userId":"6a3ddc490146560bbf360f","userName":"Raphael Chan"},"version":5}]}