メインコンテンツまでスキップ

AI Models & Assistants

すべてのレベル

Claude is our core — but the same skills travel. This section widens the lens to the whole AI world: the major assistants, how they differ, when to use which, and the techniques that transfer across all of them.

What you'll learn
  • Pick the right model for a job without re-learning the field every time
  • Move between assistants — ChatGPT, Gemini, Grok, open models — without losing your technique
  • Run capable models on your own machine, and know when that is worth it
  • Combine Claude with local models and agent frameworks

Start here

If you read one page in this section, read this one. Everything below is a branch off it.

Then follow whichever thread matches what you are doing.

The current flagships

The two September 2026 releases at the top of both stacks, read side by side.

Coming from another assistant

You already know one tool and want your habits to carry over.

Comparing before you commit

Running models yourself

Local models trade some capability for privacy, cost control, and working offline. These pages cover when that trade is worth making.

Agents beyond Claude

Everything in this section

The full list, including pages added since this index was written.

GPT-5.6 2026年8月アップデート:エフォート・スライダー、無料版のThinkボタン、そして272Kの価格クリフ

2026年8月6日、OpenAIはGPT-5.6のミッドサイクル・アップデートを出荷した。多くの記事は『Solが賢くなった』の一言で片付けたが、本当の話はもっと具体的だ。リクエストの単価を左右する6段階のエフォート制御、無料ユーザー向けのThinkボタン、コール全体を再課金する272Kトークンの価格クリフ、そしてSolでのみ動く新しいFastモード。それぞれがClaudeユーザーに何を意味するか。

Apple Foundation Models 3 と Claude ユーザーのためのオンデバイス LLM スタック

Apple の第3世代 Foundation Models(AFM 3)が WWDC 2026 で登場。3B のオンデバイス密モデル、iPhone で動作する 20B のスパース MoE、そして独自の LLM プロバイダ(Claude を含む)を持ち込める Swift ネイティブなフレームワークが揃った。実際に出荷されたもの、目立たない仕組み、そして Claude アプリを出荷する開発者にとってオンデバイスの物語がどう変わるかを解説する。

Muse Code + Muse Spark 1.2: Meta のターミナル・コーディングエージェント

2026 年 8 月 5 日、Meta はターミナル・コーディングエージェントの Muse Code と、それと共に共同学習されたモデル Muse Spark 1.2 を出荷した。追記専用イベントログが実際に何をもたらすのか、なぜ contributor ティアが標準ティアより 12.5 倍安いのか、Terminal-Bench 2.1 で Claude Code とどう並ぶのか、そしてどのエージェントを使うにせよ盗む価値のある同梱スキル 3 つ。

ネイティブなマルチエージェントAPI:OpenAIのResponsesマルチエージェントベータ vs 自作

2026年7月9日、OpenAIはResponses APIのマルチエージェントベータとSol Ultra Modeを出荷した — フロンティアベンダーで初めて『モデルが1回のHTTPリクエスト内でN個のサブエージェントを生成・統合する』ことをファーストクラスのプリミティブにした。実際に何が出荷されたか、APIの形、コスト計算、Anthropicが代わりに提供するもの、そしてネイティブプリミティブがDIYのファンアウトに勝つ場面。

Qwen3.8-Max:初のオープンウェイトMaxクラスモデル(重みが公開されたら)

2026年8月3日、AlibabaのQwen3.8-MaxはAPI専用で公開された。2.4TパラメータのMoE(アクティブ95B)、1Mコンテキスト、OpenAI/Anthropic両仕様のAPI、OSWorldでGPT-5.6 Sol Maxを上回るベンチマーク。何が本当に新しいのか、Hugging Faceのページが空のときの『オープンウェイト』とは何を意味するのか、そしてセルフホスティングを高くつかせるスパース性の計算。

Kimi K3をローカルで動かす:vLLM、DSpark、そして本当のハードウェア請求書

2026年、Moonshotの2.8Tオープンウェイト Kimi K3 をセルフホストするための決定版の実践ガイド。DSpark投機的デコードが実際に何をするか、最低限必要なハードウェア(1× DGX B300 または 16× DGX Spark)、正確なvLLMサーブコマンド、誰も警告しないプリフィル/デコードの非対称性、そしてホスト型推論が自分のラックより安くなるのはいつか。

Grok Build の内側:オープンなエージェントハーネスを読む

xAI は自社のコーディングエージェントの Rust ソース全体を Apache 2.0 で公開した — エージェントループ、ツール層、TUI、拡張システム。クレート構成がエージェント構築について教えてくれること、ここでの「オープンソース」がなぜ source-available を意味するのか、そしてツリーに残るデータコレクターのコードが物語ること。

Poolside Laguna:クラスを超える性能を発揮するオープンウェイトのコーディングモデル

Poolside AI の Laguna ファミリー(XS 2.1、S 2.1、M.1)——アクティブパラメータ 3B〜23B、100 万トークンのコンテキスト、OpenMDW-1.1 ライセンスを持つ MoE コーディングモデル。8B アクティブの S 2.1 が実際に何を打ち破るのか、3B アクティブの XS 2.1 が MacBook で動く理由、初めてとされる FP8 での RL 学習パイプライン、そして Claude、GLM-5.2、Kimi K3 より Laguna を選ぶべきタイミング。

MiniMax H3 (Hailuo 3.0): オープンウェイトのオムニ動画モデル

MiniMax は 33B の単一ストリーム Transformer をオープンソース化し、ネイティブなステレオ音声付きの 4〜15 秒の 2K 動画を 1 パスで生成します。ライセンスは US/EU/UK/韓国での自前ホスティングを静かにブロックし、オープンなベースは 768p (2K ではない) で、33B パラメータのうち約 13B は推論時にスキップできる AdaLN 分岐です。H3 が実際に何なのか、コストはいくらか、そしてクローズドな動画 API を打ち負かすのはいつなのか。

AIエージェント向け共有ログインブラウザ:ego lite パターン

新しいクラスのブラウザ — その最高の代表である ego lite は 2026年7月24日に GitHub Trending で 1 位を獲得した — が、Claude Code、Codex、Cursor、その他あらゆるコーディングエージェントに、あなたの実際のログインを継承する隔離された Chromium セッションを駆動させる。Spaces とセマンティックスナップショットが実際に解決するもの、エージェント作業で Playwright を下回る理由、自律プロセスとセッションクッキーを共有するセキュリティコスト。

Cloudflare Kitesurf:AI エージェント(人間ではなく)のために作られた初のブラウザランタイム

2026 年 8 月 6 日、Cloudflare は Kitesurf をリリースしました。Workers 上の V8 isolate だけで動くステートレスなエージェントファーストのブラウザです。一般的なエージェントタスクで Chromium より CPU とメモリを 3〜7 倍少なく使い、その代償として実時間は約 1.7 倍かかります。Rust で書かれ、Firefox の CSS パーサー(Stylo)と Blitz レンダリングエンジンを使い、Chrome DevTools Protocol 経由で Puppeteer/Playwright とそのまま互換です。このページでは、何が変わったのか、ヘッドレス Chrome の代わりにいつ選ぶべきか、Claude Code / Codex / MCP クライアントにどう組み込むか、そして現時点でまったくできない 4 つのことを解説します。

Colibrì: Run a 744B MoE From Your SSD (and What It Really Costs You)

Colibrì is a pure-C, zero-dependency engine that runs GLM-5.2 (744B), Kimi K3 (2.8T), DeepSeek V4.1 Flash and six other frontier MoE models on a 16–32 GB machine by keeping the dense layers in RAM and streaming the routed experts from NVMe. This page explains the memory hierarchy, the numbers that decide whether it is usable for you (0.05 to 6.8 tok/s), the SSD-wear and speculative-decoding gotchas, and the exact commands to get a first token.

OpenAI Agents API vs Claude Managed Agents: The Hosted Agent Loop, Compared

On September 10, 2026 OpenAI put the Codex harness behind a single API call: the Agents API public beta — durable sessions, hosted or self-hosted sandboxes, automatic compaction, tool search, subagents. Anthropic has run the same shape since April as Managed Agents. Both use the same four primitives; they differ on permissions, secrets, budgets, network policy and what the sandbox costs. The side-by-side, the request shapes, the six gotchas, and when to pick which.

Next