AI 모델 지형도: Claude, ChatGPT, Gemini 및 오픈 모델 중에서 고르기
작업에 맞는 AI 모델을 고르기 위한 오래 통하는 프레임워크 — Claude, ChatGPT, Gemini, Llama, Mistral, Grok 및 오픈/로컬 모델 전반에 걸쳐 — 그리고 이 모든 모델에 공통으로 전이되는 기술.
Claude Fable 5 & Mythos 5: 플래그십 필드 가이드
Fable 5는 Anthropic의 가장 유능한 모델입니다. API는 인-밴드로 거부(HTTP 200, stop_reason: refusal)하며, 적응형 thinking만 있고, 원시 thinking을 반환하지 않습니다. 최상위 모델을 고정한다면 통합을 작성하는 방식이 바뀝니다. Opus 4.8 대비 언제 사용할지, fallback 패턴, 프롬프팅 전환.
Claude Opus 5: The Field Guide
Opus 5 shipped 24 July 2026 as Anthropic's fourth model in two months — same $5/$25 pricing as Opus 4.8 but more than double the Frontier-Bench score, 3× ARC-AGI-3 vs the next-best model, and a new xhigh/max effort tier. The benchmarks that matter, the two 400-errors that break naive migrations, and when Opus 5 replaces Fable 5 in your stack.
Claude Sonnet 5: The Field Guide
Sonnet 5 is Claude Code's new default and Sonnet 4.6's drop-in successor — but three API constraints will 400 you on migration and a new tokenizer produces ~30% more tokens for the same text. The pricing math, the effort re-mapping (Sonnet 5 medium ≈ Sonnet 4.6 high), and the prompting shifts that catch teams out.
Claude 사용자를 위한 ChatGPT
Claude에 능숙하지만 ChatGPT를 써야 하나요? 중요한 차이점, 그대로 옮겨지는 것, 그리고 각각 언제 손을 뻗어야 하는지.
GPT-5.6 & ChatGPT Work for Claude Users
OpenAI shipped the GPT-5.6 family (Sol, Terra, Luna) and ChatGPT Work on 9 July 2026. What actually changed for a Claude user — 1.05M context, a $1 tier, an agent that runs for hours, and the parts of the launch nobody is writing about.
OpenAI Astra: The Preview Field Note
OpenAI's 'next major model' Astra was previewed on 1 August 2026 with ten Lean-verified math proofs and almost no product specs. What is actually confirmed, what is still unknown, and whether a Claude user should wait — the honest, sourced take.
Sign in with ChatGPT
OpenAI's new identity layer, three days old. What data actually crosses the line, what enterprise admins can do about it, and how to accept it in your app without giving OpenAI a foothold you didn't intend.
Ollama로 AI 모델을 로컬에서 실행하기
오픈 웨이트 모델(Llama, Mistral, Qwen…)을 여러분의 컴퓨터에서 실행하세요 — 비공개, 오프라인, 무료로 — Ollama와 함께. 설치, CLI, 그리고 코드에서 호출하기.
코딩을 위한 Claude vs GPT vs Gemini
코딩용 프런티어 모델을 고르기 위한 프레임워크(순위표가 아님) — 중요한 요소들, 그리고 왜 당신의 저장소에서 돌린 평가가 모든 벤치마크를 이기는지.
코딩 에이전트 CLI 비교
Claude Code, Codex CLI, Gemini CLI 그리고 오픈소스 터미널 에이전트 — 모델뿐 아니라 하네스(harness)가 결과를 결정하는 방식, 그리고 AGENTS.md로 설정을 이식 가능하게 유지하는 법.
Supabase Evals — Real-Backend Agent Benchmark
Supabase open-sourced a benchmark that runs Claude Code, Codex and OpenCode against real containerized Supabase stacks — not mocks. Here's what it measures, what it found about how models actually use docs and skills, and how to run it locally in an evening.
Claude 사용자를 위한 Gemini
Claude에 익숙한데 구글의 Gemini가 필요하신가요? 중요한 차이점들 — Gems, Workspace 통합, 거대한 컨텍스트, 멀티모달 — 그리고 무엇이 그대로 이어지는지.
Gemini 3.6 Flash, Flash-Lite & Flash Cyber for Claude Users
On 21 July 2026 Google skipped Gemini 3.5 Pro entirely and shipped three Flash-tier models instead. Real prices, real benchmarks, the migration gotchas, and the parts most write-ups miss — including a cybersecurity fine-tune that beat Claude Opus 4.6 at V8 bug-hunting.
Claude 사용자를 위한 Grok
Claude에 능숙한데 xAI의 Grok이 필요한가요? 중요한 차이점들 — 실시간 X/웹 검색, OpenAI 호환 API, Grok Build 코딩 에이전트 — 그리고 그대로 옮겨지는 모든 것.
Muse Code + Muse Spark 1.2: Meta's Terminal Coding Agent
On 5 August 2026 Meta shipped Muse Code, a terminal coding agent, and Muse Spark 1.2, the model co-trained with it. What the append-only event log actually buys you, why the contributor tier is 12.5× cheaper than standard, how it stacks up against Claude Code on Terminal-Bench 2.1, and the three bundled skills that are worth stealing regardless of which agent you use.
오픈소스 AI 에이전트 프레임워크
LLM 에이전트를 구축하기 위한 주요 오픈 프레임워크 — LangGraph, LlamaIndex, AutoGen, CrewAI 그리고 최소 루프 접근 방식 — 에 대한 공급자 중립적인 지도와 선택 방법.
Native Multi-Agent APIs: OpenAI's Responses Multi-Agent Beta vs Building It Yourself
On 9 July 2026 OpenAI shipped the Responses API multi-agent beta and Sol Ultra Mode — the first frontier vendor to make 'the model spawns and synthesises N subagents in one HTTP request' a first-class primitive. What actually ships, the API shape, the cost math, what Anthropic does instead, and when the native primitive beats a DIY fan-out.
A2A: 에이전트 간 프로토콜
MCP는 에이전트를 도구에 연결합니다. A2A는 에이전트를 다른 에이전트에 연결합니다 — 프레임워크, 클라우드, 회사를 넘나들며. Agent Card, task, 여덟 가지 라이프사이클 상태, 스트리밍 vs. 웹훅, 그리고 A2A가 MCP와 어떻게 결합되는지 배우세요.
모델 간 프롬프트 이식하기
Claude, GPT, Gemini 및 오픈 모델 사이에서 프롬프트를 옮기는 법 — 깔끔하게 이전되는 부분, 모델별로 조정할 부분, 그리고 마이그레이션 워크플로.
DeepSeek, Qwen 그리고 오픈 웨이트의 물결
강력한 오픈 웨이트 모델(DeepSeek, Qwen, Llama, Mistral)이 격차를 크게 좁혔습니다. '오픈 웨이트'가 무엇을 의미하는지, 폐쇄형 프런티어 대비 트레이드오프, 그리고 이를 활용하는 방법.
Claude 사용자를 위한 Kimi K2
Claude에는 익숙한데 Moonshot의 Kimi K2가 궁금하신가요? Claude Code를 실제로 겨눌 수 있는 오픈 웨이트 에이전트형 모델 — Moonshot이 Anthropic 호환 엔드포인트를 제공하기 때문입니다. 진짜로 다른 점은 무엇이고, 무엇이 그대로 넘어오는지.
GLM-5.2: Open-Weight Frontier Coding Model
Z.ai's 753B MoE (~40B active), 1M-token context, MIT-licensed — closes to within a point of Claude Opus 4.8 on long-horizon coding, at roughly one-sixth the API cost. What IndexShare actually does, how to run it locally, and where it beats Claude Code in the wild.
Inkling: Thinking Machines의 오픈 웨이트 모델
Mira Murati의 Thinking Machines Lab이 Inkling을 출시했다 — 연속적인 사고 노력(thinking-effort) 다이얼을 갖춘 9750억 파라미터 Apache 2.0 멀티모달 MoE. 출시 보도가 무엇을 놓쳤는지, 왜 의도적으로 벤치마크에서 지는지, 그리고 그 노력 노브가 실제로 무엇을 하는지.
Kimi K3: World's Largest Open-Weight Model
Moonshot AI's 2.8-trillion-parameter open-weight model beats Claude Opus 4.8 on many benchmarks and tops the Frontend Code Arena. What makes it architecturally different, when the cache-hit pricing flips the cost math, and whether it's worth reaching for.
Running Kimi K3 Locally: vLLM, DSpark & the Real Hardware Bill
The definitive practical guide to self-hosting Moonshot's 2.8T open-weight Kimi K3 in 2026. What DSpark speculative decoding actually does, minimum hardware (1× DGX B300 or 16× DGX Spark), the exact vLLM serve commands, the prefill/decode asymmetry nobody warns you about, and when hosted inference undercuts your own rack.
Grok Build 내부 들여다보기: 오픈 에이전트 하니스 읽기
xAI가 자사 코딩 에이전트의 전체 Rust 소스를 Apache 2.0으로 공개했다 — 에이전트 루프, 툴 레이어, TUI, 확장 시스템까지. 크레이트 맵이 에이전트 구축에 대해 가르쳐 주는 것, 여기서 '오픈 소스'가 왜 소스 공개(source-available)를 의미하는지, 그리고 트리에 여전히 남아 있는 데이터 수집기 코드가 무엇을 말해 주는지.
Poolside Laguna: Open-Weight Coding Models That Punch Above Their Class
Poolside AI's Laguna family (XS 2.1, S 2.1, M.1) — MoE coding models with 3B–23B active params, 1M-token context, OpenMDW-1.1 license. What the 8B-active S 2.1 actually beats, why the 3B-active XS 2.1 runs on a MacBook, the first-known RL-in-FP8 training pipeline, and when to reach for it over Claude, GLM-5.2 or Kimi K3.
생성형 미디어: 이미지, 오디오 및 비디오 AI
텍스트 채팅을 넘어서 — 이미지, 비디오, 음성, 음악을 위한 AI 지도: 주요 도구, 지속되는 기술, 그리고 중요한 권리/윤리.
AI의 실제 비용 (제공업체별 비교)
제공업체 간 AI 비용을 비교하기 위한 견고한 프레임워크 — 가격 책정 원형, 워크로드 추정 방법, 청구액을 줄이는 레버, 그리고 오픈/로컬이 유리한 시점.
AI 에이전트가 토큰을 태우는 이유 (그리고 청구서를 억제하는 법)
에이전트 실행은 채팅보다 50~1000배 비싸다 — 모델이 멍청해져서가 아니라 컨텍스트 눈덩이 때문이다. 그 메커니즘, 놀라운 숫자들, 그리고 Claude, GPT, Gemini, 오픈 모델 전반에서 실제로 청구서를 줄이는 네 가지 레버.
로컬 AI 에이전트 구축하기
로컬 오픈 웨이트 모델로 완전히 자신의 머신에서 실행되는 자율 에이전트 — 프라이빗하고, 오프라인이며, 실행 비용이 무료입니다. 아키텍처, 트레이드오프, 그리고 시작하는 방법.
Claude + 로컬 모델: 하이브리드 패턴
Claude와 로컬 오픈웨이트 모델을 시너지로 함께 작동시키세요 — 라우터, 초안 후 다듬기, 프라이버시 마스킹, 대량 전처리 — 프런티어의 추론 능력을 로컬의 프라이버시와 비용으로.
AI 게이트웨이: LiteLLM, OpenRouter, Portkey, Vercel
프로덕션 AI 게이트웨이는 여러분의 앱과 모든 모델 — Claude, GPT, Gemini, Llama — 사이의 잃어버린 라우터입니다. LiteLLM, OpenRouter, Portkey, Vercel AI Gateway를 비교하고, Claude Code를 자체 프록시를 통해 배선하는 법.
로컬 코딩 에이전트 (그리고 Claude와의 조합 방법)
당신의 머신에서 실행되며 레포지토리를 편집하는 코딩 에이전트 — Aider, Cline, Continue, OpenHands, Claude Code — 품질을 위한 Claude 또는 프라이버시를 위한 로컬 모델로 구동됩니다.
MCP로 Claude를 로컬 도구 및 에이전트에 연결하기
Model Context Protocol을 사용하면 Claude가 여러분의 컴퓨터에서 비공개로 실행되는 도구, 데이터, 에이전트를 오케스트레이션할 수 있습니다 — Claude를 두뇌로, 로컬 기능을 손으로 삼는 시너지입니다. 로컬 MCP 서버를 추가하거나 직접 만들어 보세요.
프라이빗 로컬 AI 스택 구축하기 (엔드투엔드)
모든 것을 하나로 연결하기: 로컬 모델(Ollama) + 모델 비종속 에이전트 + 로컬 MCP 서버를 통한 도구 — 당신의 기기에서 동작하는 프라이빗 어시스턴트, Claude는 선택적 스마트 레이어로.
로컬 에이전트 vs Claude: 의사결정 가이드
완전 로컬 에이전트, Claude 기반, 아니면 하이브리드? 이를 결정하는 요인들 — 프라이버시, 작업 난이도, 비용, 지연 시간, 신뢰성 — 그리고 왜 하이브리드가 종종 이기는지.
로컬 및 하이브리드 에이전트 보안
파일을 편집하고 명령을 실행할 수 있는 에이전트는 강력하면서도 위험합니다. 최소 권한, 샌드박싱, 사람의 승인, 프롬프트 인젝션 방어, 예산 상한 — 에이전트를 안전하게 실행하는 방법.
추론 모델 비교
Claude, GPT, Gemini, DeepSeek, Qwen 전반의 사고 예산(thinking budget)과 추론 강도(reasoning effort) — 언제 사고에 토큰을 쓰고, 언제 쓰지 말아야 하는가.
전이중 음성 AI: 음성 에이전트가 갑자기 진짜가 된 이유
GPT-Live와 새로운 음성 네이티브 모델들은 말하면서 동시에 듣습니다. 전이중이 무엇을 바꾸는지, 어떻게 동작하는지, 그리고 지금의 음성 AI 지형.
Claude 음성과 Opus & Sonnet (2026년 7월): 추론 우선 음성 vs GPT-Live
Anthropic이 2026년 7월 23일 Claude 음성을 Opus, Sonnet, 그리고 연결 앱으로 개방했습니다 — 하지만 턴 기반 스택을 유지했습니다. 실제로 무엇이 바뀌었는지, 통화 도중 언제 모델을 전환할지, 그리고 이 설계 베팅이 OpenAI의 풀 듀플렉스 GPT-Live와 어떻게 다른지.
토큰 속도: AI 추론이 갑자기 10-15배 빨라진 이유
웨이퍼 스케일 칩과 LPU가 프런티어 모델을 초당 수백 토큰으로 서빙합니다. 추론 속도를 실제로 제한하는 것, 새로운 하드웨어 경쟁, 그리고 빠른 토큰이 에이전트·음성·추론을 바꾸는 이유.
컴퓨터 사용 에이전트 비교
Claude, GPT, Gemini 모두 이제 화면을 조작할 수 있습니다 — 하지만 각자 다른 액션 공간, 좌표 계약, 안전 게이트를 노출합니다. 실제로 무엇이 깨지고, 그것을 견디는 루프를 어떻게 만드는가.
AI 에이전트를 위한 공유 로그인 브라우저: ego lite 패턴
새로운 부류의 브라우저 — 2026년 7월 24일 GitHub Trending #1에 오른 ego lite가 최고의 예 — 는 Claude Code, Codex, Cursor, 그리고 어떤 코딩 에이전트든 여러분의 실제 로그인을 상속받는 격리된 Chromium 세션을 몰이하게 합니다. Spaces와 semantic snapshot이 실제로 해결하는 것, 왜 에이전트 작업에서 Playwright를 앞지르는지, 그리고 자율 프로세스와 세션 쿠키를 공유하는 보안 비용.
SKILL.md: The Cross-Agent Open Standard
One SKILL.md file, thirty-two agents. What the Agent Skills open standard actually is, how progressive disclosure keeps context cheap, what breaks portability, and how to write a skill that runs unchanged in Claude Code, Codex, Gemini CLI and Cursor.