Trade places with Claude depending on the benchmark; all worth having open in a second tab.
CodingThe reference point for this page. Opus 5 leads or ties on most agentic multi-file coding benchmarks, and Claude Code is a strong terminal agent. This is the baseline the others are measured against.
ResearchExcellent at synthesis, drafting, and careful reasoning over material you provide; weaker at exhaustive literature search than the specialist tools further down this page.
NotesIncluded here for completeness. Everything below is an alternative to, or complement of, Claude — not necessarily a replacement.
Coding vs ClaudeEssentially at parity. Sol tier and Claude Opus 5 swap the lead across SWE-bench variants; margins are within noise.
Research vs ClaudeDeep Research mode is mature and produces well-structured sourced reports. Broadest plugin/connector ecosystem.
NotesBest general fallback when Claude is rate-limited.
Coding vs ClaudeSlightly behind Claude on agentic multi-file work, but the very large context window makes whole-repo and whole-codebase reading easier.
Research vs ClaudeCurrently leads several pure reasoning benchmarks. Deep Research plus NotebookLM and Workspace integration.
NotesStrongest choice when the input is a huge document set rather than a repo.
Coding vs ClaudeBehind the frontier on hard agentic coding, but the most widely used open-weight family in the West — huge tooling, fine-tune, and deployment ecosystem. Community license (open weights).
Research vs ClaudeCapable general reasoner; no research-specific features. Strong as the base model behind your own RAG or local pipeline.
NotesThe leading US open-weight option — runs locally under Ollama, so no data leaves your hardware. The Western counterpart to GLM / Qwen / DeepSeek.
Coding vs ClaudeWell below the frontier on hard agentic coding — these are small open-weight models (from ~1B up to 27B), not Gemini. Good for light coding and on-device use; the larger sizes run on a single GPU or a recent Mac.
Research vs ClaudeCapable for its size and strong multilingually; no research-specific features. Useful as a small local base model, not as a replacement for a frontier reasoner.
NotesGoogle’s open-weight family — the lightweight counterpart to Gemini, in the same class as Llama and Qwen. Open weights under Google’s Gemma terms; runs locally under Ollama.
Coding vs ClaudeRuns on OpenAI frontier models, so raw capability is close to GPT-5.6; the value is deep integration with Windows and the Microsoft 365 apps rather than a coding edge over Claude.
Research vs ClaudeGrounded web answers with citations, plus access to your own Word, Excel, and Outlook content in the enterprise version. Weaker at extended reasoning than Claude.
NotesThe path of least resistance if your work already lives in Microsoft 365. Enterprise (M365 Copilot) keeps prompts inside your tenant; the free consumer version does not.
Coding vs ClaudeClose to frontier and noticeably cheaper; reportedly completes long agentic runs in fewer turns.
Research vs ClaudeFast on live web and social data. Weaker sourcing discipline — verify citations.
NotesValue play for high-volume coding. Moderation looser than institutional work may want.