본문으로 건너뛰기

AI Models & Assistants

모든 레벨

Claude is our core — but the same skills travel. This section widens the lens to the whole AI world: the major assistants, how they differ, when to use which, and the techniques that transfer across all of them.

What you'll learn
  • Pick the right model for a job without re-learning the field every time
  • Move between assistants — ChatGPT, Gemini, Grok, open models — without losing your technique
  • Run capable models on your own machine, and know when that is worth it
  • Combine Claude with local models and agent frameworks

Start here

If you read one page in this section, read this one. Everything below is a branch off it.

Then follow whichever thread matches what you are doing.

The current flagships

The two September 2026 releases at the top of both stacks, read side by side.

Coming from another assistant

You already know one tool and want your habits to carry over.

Comparing before you commit

Running models yourself

Local models trade some capability for privacy, cost control, and working offline. These pages cover when that trade is worth making.

Agents beyond Claude

Everything in this section

The full list, including pages added since this index was written.

Claude Opus 5: 실전 가이드

Opus 5는 2026년 7월 24일에 출시되었으며, 두 달 만에 나온 Anthropic의 네 번째 모델입니다. Opus 4.8과 동일한 $5/$25 가격이지만 Frontier-Bench 점수는 두 배 이상, ARC-AGI-3는 다음으로 좋은 모델의 3배이며 새로운 xhigh/max effort 계층이 추가되었습니다. 중요한 벤치마크, 순진한 마이그레이션을 깨뜨리는 두 가지 400 오류, 그리고 Opus 5가 스택에서 Fable 5를 대체하는 시점을 다룹니다.

GPT-5.6 2026년 8월 업데이트: 노력 슬라이더, 무료 Think 버튼 & 272K 가격 절벽

2026년 8월 6일 OpenAI가 GPT-5.6에 대한 사이클 중간 업데이트를 배포했고, 대부분의 리뷰는 이를 'Sol이 더 똑똑해졌다'로 축약했다. 진짜 이야기: 요청 가격 방식을 바꾸는 6단계 노력 제어, 무료 계층 Think 버튼, 전체 호출을 다시 가격 매기는 272K 토큰 가격 절벽, Sol에서만 작동하는 Fast 모드. Claude 사용자에게 각각 의미하는 바.

Claude 사용자를 위한 Apple Foundation Models 3와 온디바이스 LLM 스택

Apple의 3세대 Foundation Models(AFM 3)가 WWDC 2026에서 등장했습니다 — 3B 밀집 온디바이스 모델, iPhone에서 실행되는 20B 스파스 MoE, 그리고 이제 여러분의 LLM 프로바이더(Claude 포함)를 가져올 수 있게 하는 Swift 네이티브 프레임워크. 실제로 출시된 것, 명백하지 않은 메커니즘, 그리고 Claude 앱을 배포하는 사람에게 온디바이스 스토리를 어떻게 바꾸는지.

Muse Code + Muse Spark 1.2: Meta의 터미널 코딩 에이전트

2026년 8월 5일, Meta는 터미널 코딩 에이전트인 Muse Code와 함께 co-training된 모델 Muse Spark 1.2를 출시했습니다. append-only 이벤트 로그가 실제로 무엇을 제공하는지, contributor 티어가 표준 대비 왜 12.5배 저렴한지, Terminal-Bench 2.1에서 Claude Code와 어떻게 비교되는지, 그리고 어떤 에이전트를 쓰든 훔쳐올 만한 세 가지 번들 스킬은 무엇인지.

네이티브 멀티 에이전트 API: OpenAI의 Responses 멀티 에이전트 베타 vs 직접 구축하기

2026년 7월 9일, OpenAI는 Responses API 멀티 에이전트 베타와 Sol Ultra Mode를 출시했습니다. '모델이 하나의 HTTP 요청으로 N개의 서브에이전트를 스폰하고 종합하는' 것을 최상위 프리미티브로 만든 최초의 프런티어 벤더입니다. 실제로 출시된 것, API 형태, 비용 계산, Anthropic의 대안, 그리고 네이티브 프리미티브가 DIY 팬아웃을 이기는 경우를 다룹니다.

Qwen3.8-Max: 최초의 오픈 웨이트 Max급 모델 (가중치가 실제로 공개된다면)

2026년 8월 3일: 알리바바의 Qwen3.8-Max는 2.4T 파라미터 MoE(95B 활성), 1M 컨텍스트, OpenAI/Anthropic 규격 이중 API를 갖추고 API 전용으로 출시되었으며, OSWorld에서 GPT-5.6 Sol Max를 앞서는 벤치마크 점수를 기록했다. 진정으로 새로운 점, Hugging Face 페이지가 비어 있을 때 '오픈 웨이트'가 실제로 의미하는 바, 그리고 자체 호스팅을 비싸게 만드는 희소성 수학.

MiniMax H3 (Hailuo 3.0): 오픈 웨이트 옴니 비디오 모델

MiniMax는 33B 단일 스트림 트랜스포머를 오픈소스로 공개했습니다. 이 모델은 한 번의 패스로 네이티브 스테레오 오디오와 함께 2K 해상도의 4~15초 비디오를 생성합니다. 라이선스는 미국/EU/영국/한국에서의 셀프 호스팅을 조용히 차단하며, 오픈 베이스는 2K가 아닌 768p이고, 33B 파라미터 중 약 13B는 추론 시 건너뛸 수 있는 AdaLN 브랜치입니다. H3가 실제로 무엇이며, 비용은 얼마이고, 언제 클로즈드 비디오 API를 이길 수 있는지.

AI 에이전트를 위한 공유 로그인 브라우저: ego lite 패턴

새로운 부류의 브라우저 — 2026년 7월 24일 GitHub Trending #1에 오른 ego lite가 최고의 예 — 는 Claude Code, Codex, Cursor, 그리고 어떤 코딩 에이전트든 여러분의 실제 로그인을 상속받는 격리된 Chromium 세션을 몰이하게 합니다. Spaces와 semantic snapshot이 실제로 해결하는 것, 왜 에이전트 작업에서 Playwright를 앞지르는지, 그리고 자율 프로세스와 세션 쿠키를 공유하는 보안 비용.

Cloudflare Kitesurf: The First Browser Runtime Built for AI Agents (Not Humans)

On August 6, 2026 Cloudflare launched Kitesurf, a stateless, agent-first browser that runs entirely in V8 isolates on Workers. It uses 3-7x less CPU and memory than Chromium on common agent tasks — at the cost of ~1.7x wall-clock time. Written in Rust, using Firefox's CSS parser (Stylo) and the Blitz rendering engine, drop-in compatible with Puppeteer/Playwright via the Chrome DevTools Protocol. This page: what changed, when to reach for it instead of headless Chrome, how to wire it into Claude Code / Codex / your MCP client, and the four things it flat-out can't do yet.

Colibrì: Run a 744B MoE From Your SSD (and What It Really Costs You)

Colibrì is a pure-C, zero-dependency engine that runs GLM-5.2 (744B), Kimi K3 (2.8T), DeepSeek V4.1 Flash and six other frontier MoE models on a 16–32 GB machine by keeping the dense layers in RAM and streaming the routed experts from NVMe. This page explains the memory hierarchy, the numbers that decide whether it is usable for you (0.05 to 6.8 tok/s), the SSD-wear and speculative-decoding gotchas, and the exact commands to get a first token.

OpenAI Agents API vs Claude Managed Agents: The Hosted Agent Loop, Compared

On September 10, 2026 OpenAI put the Codex harness behind a single API call: the Agents API public beta — durable sessions, hosted or self-hosted sandboxes, automatic compaction, tool search, subagents. Anthropic has run the same shape since April as Managed Agents. Both use the same four primitives; they differ on permissions, secrets, budgets, network policy and what the sandbox costs. The side-by-side, the request shapes, the six gotchas, and when to pick which.

Next