El panorama de modelos de IA: elegir entre Claude, ChatGPT, Gemini y modelos abiertos
Un marco duradero para elegir el modelo de IA adecuado para una tarea — entre Claude, ChatGPT, Gemini, Llama, Mistral, Grok y los modelos abiertos/locales — además de qué habilidades se transfieren a todos ellos.
Claude Fable 5 & Mythos 5: The Flagship Field Guide
Fable 5 is Anthropic's most capable model. Its API refuses in-band (HTTP 200, stop_reason: refusal), has adaptive thinking only, and never returns raw thinking. If you're pinning a top-tier model, this changes how you write integrations. When to use it vs Opus 4.8, the fallback pattern, the prompting shifts.
Claude Opus 5: The Field Guide
Opus 5 shipped 24 July 2026 as Anthropic's fourth model in two months — same $5/$25 pricing as Opus 4.8 but more than double the Frontier-Bench score, 3× ARC-AGI-3 vs the next-best model, and a new xhigh/max effort tier. The benchmarks that matter, the two 400-errors that break naive migrations, and when Opus 5 replaces Fable 5 in your stack.
Claude Sonnet 5: The Field Guide
Sonnet 5 is Claude Code's new default and Sonnet 4.6's drop-in successor — but three API constraints will 400 you on migration and a new tokenizer produces ~30% more tokens for the same text. The pricing math, the effort re-mapping (Sonnet 5 medium ≈ Sonnet 4.6 high), and the prompting shifts that catch teams out.
ChatGPT para usuarios de Claude
¿Dominas Claude y necesitas usar ChatGPT? Las diferencias que importan, lo que se transfiere y cuándo recurrir a cada uno.
GPT-5.6 & ChatGPT Work for Claude Users
OpenAI shipped the GPT-5.6 family (Sol, Terra, Luna) and ChatGPT Work on 9 July 2026. What actually changed for a Claude user — 1.05M context, a $1 tier, an agent that runs for hours, and the parts of the launch nobody is writing about.
OpenAI Astra: The Preview Field Note
OpenAI's 'next major model' Astra was previewed on 1 August 2026 with ten Lean-verified math proofs and almost no product specs. What is actually confirmed, what is still unknown, and whether a Claude user should wait — the honest, sourced take.
Sign in with ChatGPT
OpenAI's new identity layer, three days old. What data actually crosses the line, what enterprise admins can do about it, and how to accept it in your app without giving OpenAI a foothold you didn't intend.
Ejecutar modelos de IA localmente con Ollama
Ejecuta modelos de pesos abiertos (Llama, Mistral, Qwen…) en tu propia máquina — privado, sin conexión, gratis de ejecutar — con Ollama. Instalación, CLI y cómo llamarlo desde código.
Claude vs GPT vs Gemini para programar
Un marco de decisión (no una tabla de clasificación) para elegir un modelo de frontera para programar — los factores que importan, y por qué una evaluación sobre tu propio repositorio supera a cualquier benchmark.
CLIs de agentes de codificación comparadas
Claude Code, Codex CLI, Gemini CLI y los agentes de terminal de código abierto: cómo el harness (no solo el modelo) decide tus resultados, y cómo mantener tu configuración portable con AGENTS.md.
Supabase Evals — Real-Backend Agent Benchmark
Supabase open-sourced a benchmark that runs Claude Code, Codex and OpenCode against real containerized Supabase stacks — not mocks. Here's what it measures, what it found about how models actually use docs and skills, and how to run it locally in an evening.
Gemini para usuarios de Claude
¿Dominas Claude y necesitas Gemini de Google? Las diferencias que importan — Gems, integración con Workspace, contexto enorme, multimodal — y lo que se transfiere.
Gemini 3.6 Flash, Flash-Lite & Flash Cyber for Claude Users
On 21 July 2026 Google skipped Gemini 3.5 Pro entirely and shipped three Flash-tier models instead. Real prices, real benchmarks, the migration gotchas, and the parts most write-ups miss — including a cybersecurity fine-tune that beat Claude Opus 4.6 at V8 bug-hunting.
Grok para usuarios de Claude
¿Dominas Claude y necesitas Grok de xAI? Las diferencias que importan — búsqueda en tiempo real en X/web, una API compatible con OpenAI y el agente de programación Grok Build — más todo lo que se transfiere.
Muse Code + Muse Spark 1.2: Meta's Terminal Coding Agent
On 5 August 2026 Meta shipped Muse Code, a terminal coding agent, and Muse Spark 1.2, the model co-trained with it. What the append-only event log actually buys you, why the contributor tier is 12.5× cheaper than standard, how it stacks up against Claude Code on Terminal-Bench 2.1, and the three bundled skills that are worth stealing regardless of which agent you use.
Frameworks de código abierto para agentes de IA
Un mapa neutral respecto al proveedor de los principales frameworks abiertos para construir agentes LLM — LangGraph, LlamaIndex, AutoGen, CrewAI y el enfoque de bucle mínimo — y cómo elegir.
Native Multi-Agent APIs: OpenAI's Responses Multi-Agent Beta vs Building It Yourself
On 9 July 2026 OpenAI shipped the Responses API multi-agent beta and Sol Ultra Mode — the first frontier vendor to make 'the model spawns and synthesises N subagents in one HTTP request' a first-class primitive. What actually ships, the API shape, the cost math, what Anthropic does instead, and when the native primitive beats a DIY fan-out.
A2A: The Agent-to-Agent Protocol
MCP connects agents to tools. A2A connects agents to other agents — across frameworks, clouds, and companies. Learn the Agent Card, tasks, the eight lifecycle states, streaming vs. webhooks, and how A2A composes with MCP.
Portar Prompts entre Modelos
Mueve un prompt entre Claude, GPT, Gemini y modelos abiertos — qué se transfiere sin cambios, qué ajustar por modelo y un flujo de migración.
DeepSeek, Qwen y la ola de pesos abiertos
Modelos potentes de pesos abiertos (DeepSeek, Qwen, Llama, Mistral) cerraron gran parte de la brecha. Qué significa 'pesos abiertos', las concesiones frente a la frontera cerrada y cómo usarlos.
Kimi K2 for Claude Users
Moonshot's Kimi K2 is an open-weight, trillion-parameter agent built for long tool-use chains. What actually makes it different from Claude — native INT4 weights, 200–300 sequential tool calls, a permissive license — and when it's worth reaching for.
GLM-5.2: Open-Weight Frontier Coding Model
Z.ai's 753B MoE (~40B active), 1M-token context, MIT-licensed — closes to within a point of Claude Opus 4.8 on long-horizon coding, at roughly one-sixth the API cost. What IndexShare actually does, how to run it locally, and where it beats Claude Code in the wild.
Inkling: Thinking Machines' Open-Weights Model
Mira Murati's Thinking Machines Lab shipped Inkling — a 975B-parameter Apache 2.0 multimodal MoE with a continuous thinking-effort dial. What the launch coverage got wrong, why it deliberately loses benchmarks, and what the effort knob actually does.
Kimi K3: World's Largest Open-Weight Model
Moonshot AI's 2.8-trillion-parameter open-weight model beats Claude Opus 4.8 on many benchmarks and tops the Frontend Code Arena. What makes it architecturally different, when the cache-hit pricing flips the cost math, and whether it's worth reaching for.
Running Kimi K3 Locally: vLLM, DSpark & the Real Hardware Bill
The definitive practical guide to self-hosting Moonshot's 2.8T open-weight Kimi K3 in 2026. What DSpark speculative decoding actually does, minimum hardware (1× DGX B300 or 16× DGX Spark), the exact vLLM serve commands, the prefill/decode asymmetry nobody warns you about, and when hosted inference undercuts your own rack.
Inside Grok Build: Reading an Open Agent Harness
xAI published the full Rust source of its coding agent under Apache 2.0 — the agent loop, tool layer, TUI and extension system. What the crate map teaches you about building agents, why 'open source' here means source-available, and what the data-collector code still in the tree tells you.
Poolside Laguna: Open-Weight Coding Models That Punch Above Their Class
Poolside AI's Laguna family (XS 2.1, S 2.1, M.1) — MoE coding models with 3B–23B active params, 1M-token context, OpenMDW-1.1 license. What the 8B-active S 2.1 actually beats, why the 3B-active XS 2.1 runs on a MacBook, the first-known RL-in-FP8 training pipeline, and when to reach for it over Claude, GLM-5.2 or Kimi K3.
Medios generativos: IA de imagen, audio y vídeo
Más allá del chat de texto — un mapa de la IA para imágenes, vídeo, voz y música: las herramientas principales, las habilidades duraderas y los derechos/ética que importan.
Cuánto cuesta realmente la IA (entre proveedores)
Un marco duradero para comparar el coste de la IA entre proveedores: los arquetipos de precios, cómo estimar una carga de trabajo, las palancas para reducir la factura y cuándo gana lo abierto/local.
Why AI Agents Burn Tokens (and How to Cap the Bill)
Agentic runs cost 50–1000x a chat — not because the model is dumber, but because of the context snowball. The mechanism, the surprising numbers, and the four levers that actually cut the bill across Claude, GPT, Gemini, and open models.
Crear agentes de IA locales
Agentes autónomos que se ejecutan por completo en tu máquina con un modelo local de pesos abiertos: privados, sin conexión y gratis de ejecutar. La arquitectura, las concesiones y cómo empezar.
Claude + modelos locales: patrones híbridos
Haz que Claude y los modelos locales de pesos abiertos trabajen en sinergia — router, borrador-y-refinamiento, redacción para privacidad, preprocesamiento masivo — para lograr el razonamiento de frontera con la privacidad y el coste de lo local.
AI Gateways: LiteLLM, OpenRouter, Portkey, Vercel
A production AI gateway is the missing router between your app and every model — Claude, GPT, Gemini, Llama. Compare LiteLLM, OpenRouter, Portkey, and Vercel AI Gateway, then wire Claude Code through your own proxy.
Agentes de codificación locales (y cómo se combinan con Claude)
Agentes de codificación que se ejecutan en tu máquina y editan tu repositorio — Aider, Cline, Continue, OpenHands, Claude Code — impulsados por Claude para la calidad o por un modelo local para la privacidad.
Conecta Claude a herramientas y agentes locales con MCP
El Model Context Protocol permite a Claude orquestar herramientas, datos y agentes que se ejecutan de forma privada en tu propia máquina — la sinergia Claude-como-cerebro, capacidades-locales. Añade o construye un servidor MCP local.
Construye un stack de IA local y privado (de principio a fin)
Une todas las piezas: un modelo local (Ollama) + un agente agnóstico al modelo + herramientas vía un servidor MCP local — un asistente privado en tu propia máquina, con Claude como capa inteligente opcional.
¿Agente local o Claude? Una guía de decisión
¿Agente totalmente local, con Claude, o híbrido? Los factores que lo deciden — privacidad, dificultad de la tarea, coste, latencia, fiabilidad — y por qué a menudo gana el híbrido.
Proteger agentes locales e híbridos
Un agente que puede editar archivos y ejecutar comandos es poderoso y peligroso. Privilegio mínimo, sandboxing, aprobación humana, defensa contra inyección de prompts, límites de presupuesto: cómo ejecutar agentes de forma segura.
Modelos de razonamiento comparados
Presupuestos de pensamiento y esfuerzo de razonamiento en Claude, GPT, Gemini, DeepSeek y Qwen — cuándo gastar tokens en pensar, y cuándo no.
Full-Duplex Voice AI: Why Voice Agents Suddenly Got Real
GPT-Live and the new speech-native models listen and talk at the same time. What full-duplex changes, how it works, and the voice-AI landscape now.
Claude Voice with Opus & Sonnet (July 2026): Reasoning-First Voice vs GPT-Live
Anthropic opened Claude voice to Opus, Sonnet, and connected apps on July 23, 2026 — but kept the turn-based stack. What actually changed, when to switch models mid-call, and how the design bet differs from OpenAI's full-duplex GPT-Live.
Token Speed: Why AI Inference Suddenly Got 10-15× Faster
Wafer-scale chips and LPUs are serving frontier models at hundreds of tokens per second. What actually limits inference speed, the new hardware race, and why fast tokens change agents, voice and reasoning.
Computer-Use Agents Compared
Claude, GPT and Gemini can all drive a screen now — but each exposes a different action space, coordinate contract and safety gate. What actually breaks, and how to build a loop that survives it.
Shared-Login Browsers for AI Agents: The ego lite Pattern
A new class of browser — best exemplified by ego lite, which hit #1 on GitHub Trending on July 24, 2026 — lets Claude Code, Codex, Cursor, and any other coding agent drive an isolated Chromium session that inherits your real logins. What Spaces and semantic snapshots actually solve, why it undercuts Playwright for agent work, and the security cost of sharing session cookies with an autonomous process.
SKILL.md: The Cross-Agent Open Standard
One SKILL.md file, thirty-two agents. What the Agent Skills open standard actually is, how progressive disclosure keeps context cheap, what breaks portability, and how to write a skill that runs unchanged in Claude Code, Codex, Gemini CLI and Cursor.