Anthropic · Hosted API
Price $1.00/M input · $5.00/M output
Context 200K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · anthropic/claude-haiku-4.5
Moderate estimate $35.00/mo
Best for: Summaries, extraction, rewriting, support tasks, and short Q&A.
Not ideal for: Complex reasoning or long-context synthesis.
premium Vision Tools Reasoning text writing summarization classification extraction fast-response
CompareChecking fit…
Anthropic · Hosted API
Price Input not observed · Output not observed
Context 1M Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · anthropic/claude-opus-4.8
Moderate estimate — Maximum output 128,000 tokens
Benchmark evidence · not a rank
88.6/100 SWE-bench Verified · coding · vVerified · verified 23d
Published provider ceiling, not guaranteed visible answer length; reasoning, input, endpoint, or account conditions may reduce it.
Best for: Complex reasoning, proposal work, long-form synthesis, and high-stakes review.
Not ideal for: Bulk extraction, simple summaries, or cheap high-volume routing.
premium Vision Tools Reasoning text reasoning research writing summarization long-context agentic
CompareChecking fit…
Anthropic · Hosted API
Price Input not observed · Output not observed
Context 1M Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · anthropic/claude-sonnet-4.6
Moderate estimate — Maximum output 128,000 tokens
Published provider ceiling, not guaranteed visible answer length; reasoning, input, endpoint, or account conditions may reduce it.
Best for: Strong writing, coding, document analysis, and research drafts.
Not ideal for: High-volume, cost-sensitive extraction or classification.
premium Vision Tools Reasoning text code reasoning writing research summarization
CompareChecking fit…
Anthropic · Hosted API
Price $2.00/M input · $10.00/M output
Context 1M Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · anthropic/claude-sonnet-5
Moderate estimate $70.00/mo
Best for: Everyday reasoning, writing, code, and research summaries.
Not ideal for: Ultra-low-cost bulk extraction at scale.
premium Vision Tools Reasoning text code reasoning writing research summarization
CompareChecking fit…
Google · Hosted API
Price Input not observed · Output not observed
Context 1.0M Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · google/gemini-3.1-pro-preview
Moderate estimate — Maximum output 65,536 tokens
Published provider ceiling, not guaranteed visible answer length; reasoning, input, endpoint, or account conditions may reduce it.
Best for: Long-context analysis, multimodal files, research, and document-heavy work.
Not ideal for: Short, simple chat where its context window and cost are overkill.
premium Vision Tools Reasoning text image-understanding video-understanding reasoning research long-context multimodal
CompareChecking fit…
OpenAI · Hosted API
Price Input not observed · Output not observed
Context 1.1M Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · openai/gpt-5.4
Moderate estimate — Maximum output 128,000 tokens
Published provider ceiling, not guaranteed visible answer length; reasoning, input, endpoint, or account conditions may reduce it.
Best for: Coding, reasoning, data analysis, and agent workflows.
Not ideal for: Cheap bulk tasks better handled by a budget model.
premium Vision Tools Reasoning text code reasoning data-analysis tool-use agentic multimodal
CompareChecking fit…
OpenAI · Hosted API
Price Input not observed · Output not observed
Context 1.1M Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · openai/gpt-5.5
Moderate estimate — Maximum output 128,000 tokens
Published provider ceiling, not guaranteed visible answer length; reasoning, input, endpoint, or account conditions may reduce it.
Best for: Hard reasoning, complex coding, research synthesis, and tool-heavy workflows.
Not ideal for: Simple summaries or bulk extraction where a budget model is sufficient.
premium Vision Tools Reasoning text code reasoning research tool-use agentic multimodal
CompareChecking fit…
OpenAI · Hosted API
Price Input not observed · Output not observed
Context 200K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · openai/o4-mini
Moderate estimate — Maximum output 100,000 tokens
Published provider ceiling, not guaranteed visible answer length; reasoning, input, endpoint, or account conditions may reduce it.
Best for: Math, logic, coding help, and structured planning at lower cost.
Not ideal for: Long-form writing or multimodal work.
premium Vision Tools Reasoning text code reasoning planning low-cost
CompareChecking fit…
Mistral · Hosted API
Price $0.30/M input · $0.90/M output
Context 256K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · mistralai/codestral-2508
Moderate estimate $8.70/mo
Best for: Code completion, repo tasks, and coding assistants.
Not ideal for: General-purpose chat.
standard Tools code tool-use open-weight
CompareChecking fit…
DeepSeek · Hosted API
Price $0.80/M input · $0.80/M output
Context 128K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · deepseek/deepseek-r1-distill-llama-70b
Moderate estimate $18.40/mo Parameters 70B total
Best for: Stronger self-hosted reasoning and coding than small distills.
Not ideal for: Low-memory local setups.
standard Reasoning text reasoning code open-weight self-hostable
CompareChecking fit…
Google · Hosted API
Price Input not observed · Output not observed
Context 1.0M Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · google/gemini-3-flash-preview
Moderate estimate — Maximum output 65,536 tokens
Benchmark evidence · not a rank
78/100 SWE-bench Verified · coding · vVerified · verified 23d
Published provider ceiling, not guaranteed visible answer length; reasoning, input, endpoint, or account conditions may reduce it.
Best for: Fast responses, multimodal summaries, and lightweight extraction.
Not ideal for: High-stakes reasoning or long-form synthesis.
standard Vision Tools Reasoning text image-understanding reasoning summarization extraction fast-response multimodal
CompareChecking fit…
Z.ai · Hosted API
Price $0.30/M input · $0.90/M output
Context 131K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · z-ai/glm-4.6v
Moderate estimate $8.70/mo
Best for: Vision-language, screenshots, and lightweight visual reasoning.
Not ideal for: Text-only bulk work.
standard Vision Tools Reasoning text image-understanding multimodal fast-response
CompareChecking fit…
Z.ai · Hosted API
Price $0.40/M input · $1.75/M output
Context 203K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · z-ai/glm-4.7
Moderate estimate $13.25/mo
Best for: General chat, coding help, and structured outputs.
Not ideal for: Premium reasoning or multimodal analysis.
standard Tools Reasoning text code reasoning writing low-cost
CompareChecking fit…
xAI · Hosted API
Price $1.25/M input · $2.50/M output
Context 1M Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · x-ai/grok-4.3
Moderate estimate $32.50/mo
Best for: Conversational tasks, quick answers, and topical summaries when connected to search.
Not ideal for: Deep multi-step reasoning or heavy coding.
standard Vision Tools Reasoning text reasoning writing summarization fast-response
CompareChecking fit…
Moonshot AI · Hosted API
Price $0.60/M input · $2.50/M output
Context 262K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · moonshotai/kimi-k2-thinking
Moderate estimate $19.50/mo
Best for: Lower-cost reasoning, planning, code review, and long-form answers.
Not ideal for: Multimodal tasks or latency-critical chat.
standard Tools Reasoning text code reasoning planning low-cost
CompareChecking fit…
MiniMax · Hosted API
Price Input not observed · Output not observed
Context 1M Availability Not observed Offer Ollama
More technical evidence Exact ID · minimax-m3:cloud
Moderate estimate —
Benchmark evidence · not a rank
80.5/100 SWE-bench Verified · coding · vVerified · verified 23d
Best for: Chat, writing, summaries, and simple reasoning.
Not ideal for: Hard reasoning, complex coding, or tool-heavy agents.
standard Vision Tools Reasoning text writing summarization classification low-cost
CompareChecking fit…
MiniMax · Hosted API
Price $0.30/M input · $1.20/M output
Context 1M Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · minimax/minimax-m3
Moderate estimate $9.60/mo
Benchmark evidence · not a rank
80.5/100 SWE-bench Verified · coding · vVerified · verified 23d
Best for: Chat, writing, summaries, and simple reasoning.
Not ideal for: Hard reasoning, complex coding, or tool-heavy agents.
standard Vision Tools Reasoning text writing summarization classification low-cost
CompareChecking fit…
DeepSeek · Hosted API
Price $0.27/M input · $0.40/M output
Context 131K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · deepseek/deepseek-v3.2
Moderate estimate $6.58/mo Architecture Transformer with MLA attention and mixture-of-experts feed-forward blocks
Best for: Coding, extraction, summarization, and cost-sensitive reasoning.
Not ideal for: Multimodal work or polished long-form writing.
value Tools Reasoning text code reasoning extraction summarization low-cost
CompareChecking fit…
Google · Hosted API
Price $0.08/M input · $0.45/M output
Context 131K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · google/gemma-3-27b-it
Moderate estimate $2.95/mo Parameters 27B total Architecture Dense decoder-only Transformer with interleaved local and global attention License Gemma Terms of Use
Best for: Local/simple assistant, classification, extraction, and summaries.
Not ideal for: Hard reasoning, coding agents, or long-context work.
value Vision Tools text summarization classification extraction open-weight low-cost
CompareChecking fit…
Google · Model download
Price Input not observed · Output not observed
Context 131K Availability Not observed Offer LM Studio
More technical evidence Exact ID · lmstudio-community/gemma-3-27B-it-qat-GGUF:Q4_0
Moderate estimate — Artifact GGUF · Q4_0 Reported download 15.6 GB Minimum system memory 16 GB · LM Studio Parameters 27B total Architecture Dense decoder-only Transformer with interleaved local and global attention License Gemma Terms of Use
Reported download size is storage evidence, not a RAM or VRAM requirement. The separate 16 GB · LM Studio minimum is source-backed, but still does not prove complete hardware fit.
Best for: Local/simple assistant, classification, extraction, and summaries.
Not ideal for: Hard reasoning, coding agents, or long-context work.
value Vision Tools text summarization classification extraction open-weight low-cost
CompareChecking fit…
Google · Hosted API
Price $0.07/M input · $0.34/M output
Context 262K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · google/gemma-4-26b-a4b-it
Moderate estimate $2.42/mo
Best for: Local multimodal tasks with lower active-parameter cost.
Not ideal for: Tasks needing maximum local quality.
value Vision Tools Reasoning text image-understanding multimodal open-weight fast-response local
CompareChecking fit…
Meta · Hosted API
Price $0.10/M input · $0.32/M output
Context 131K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · meta-llama/llama-3.3-70b-instruct
Moderate estimate $2.96/mo Parameters 70B total License Llama 3.3 Community License Agreement
Best for: Self-hosted general assistant, RAG, and internal tools.
Not ideal for: Frontier reasoning or multimodal tasks.
value Tools text code reasoning open-weight self-hostable low-cost
CompareChecking fit…
Meta · Model download
Price Input not observed · Output not observed
Context 131K Availability Not observed Offer LM Studio
More technical evidence Exact ID · lmstudio-community/Llama-3.3-70B-Instruct-GGUF:Q4_K_M
Moderate estimate — Artifact GGUF · Q4_K_M Reported download 42.5 GB Minimum system memory 40 GB · LM Studio Parameters 70B total License Llama 3.3 Community License Agreement
Reported download size is storage evidence, not a RAM or VRAM requirement. The separate 40 GB · LM Studio minimum is source-backed, but still does not prove complete hardware fit.
Best for: Self-hosted general assistant, RAG, and internal tools.
Not ideal for: Frontier reasoning or multimodal tasks.
value Tools text code reasoning open-weight self-hostable low-cost
CompareChecking fit…
Meta · Hosted API
Price $0.20/M input · $0.80/M output
Context 1.0M Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · meta-llama/llama-4-maverick
Moderate estimate $6.40/mo Parameters 400B total · 17B active Architecture Autoregressive mixture-of-experts model with early-fusion multimodality License Llama 4 Community License Agreement
Best for: Self-hosting, customization, and privacy-aware workloads.
Not ideal for: Managed convenience or top-tier reasoning out of the box.
value Vision Tools text code reasoning open-weight self-hostable privacy-sensitive
CompareChecking fit…
MiniMax · Hosted API
Price $0.26/M input · $1.02/M output
Context 205K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · minimax/minimax-m2
Moderate estimate $8.16/mo
Best for: Agentic coding, tool-heavy workflows, and long-context tasks.
Not ideal for: Multimodal or vision work.
value Tools Reasoning text code reasoning tool-use agentic
CompareChecking fit…
Mistral · Hosted API
Price $0.20/M input · $0.20/M output
Context 262K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · mistralai/ministral-14b-2512
Moderate estimate $4.60/mo
Best for: Better local chat and summaries than tiny models, at manageable cost.
Not ideal for: Frontier reasoning or coding.
value Vision Tools text image-understanding writing summarization local open-weight
CompareChecking fit…
Mistral · Hosted API
Price $0.15/M input · $0.15/M output
Context 262K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · mistralai/ministral-8b-2512
Moderate estimate $3.45/mo
Best for: Fast local chat, summaries, extraction, and light multimodal tasks.
Not ideal for: Complex reasoning or hard coding.
value Vision Tools text image-understanding fast-response local low-cost
CompareChecking fit…
Mistral · Hosted API
Price $0.09/M input · $0.25/M output
Context 128K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · mistralai/mistral-small-3.2-24b-instruct
Moderate estimate $2.63/mo License Apache License 2.0
Best for: Fast chat, extraction, support automation, and simple code.
Not ideal for: Complex reasoning or research synthesis.
value Vision Tools text code writing summarization classification low-cost
CompareChecking fit…
Mistral · Model download
Price Input not observed · Output not observed
Context 128K Availability Not observed Offer LM Studio
More technical evidence Exact ID · lmstudio-community/Mistral-Small-3.2-24B-Instruct-2506-GGUF:Q4_K_M
Moderate estimate — Artifact GGUF · Q4_K_M Reported download 14.3 GB Minimum system memory 14 GB · LM Studio License Apache License 2.0
Reported download size is storage evidence, not a RAM or VRAM requirement. The separate 14 GB · LM Studio minimum is source-backed, but still does not prove complete hardware fit.
Best for: Fast chat, extraction, support automation, and simple code.
Not ideal for: Complex reasoning or research synthesis.
value Vision Tools text code writing summarization classification low-cost
CompareChecking fit…
NVIDIA · Hosted API
Price $0.09/M input · $0.40/M output
Context 1M Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · nvidia/nemotron-3-super-120b-a12b
Moderate estimate $2.90/mo Architecture Mamba2-Transformer hybrid LatentMoE with multi-token prediction
Best for: Enterprise assistants, synthetic data, reasoning, and internal RAG.
Not ideal for: Latency-critical chat or multimodal analysis.
value Tools Reasoning text code reasoning summarization open-weight self-hostable
CompareChecking fit…
Alibaba · Hosted API
Price $0.12/M input · $0.80/M output
Context 262K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · qwen/qwen3-coder-next
Moderate estimate $4.80/mo
Best for: Long-horizon coding, repo work, tool recovery, and local dev agents.
Not ideal for: General chat or lightweight summarization.
value Tools code tool-use agentic planning open-weight self-hostable
CompareChecking fit…
Alibaba · Hosted API
Price $0.13/M input · $0.52/M output
Context 262K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · qwen/qwen3-vl-30b-a3b-instruct
Moderate estimate $4.16/mo
Best for: Local image/document understanding and multimodal assistants.
Not ideal for: Low-memory laptops.
value Vision Tools image-understanding multimodal open-weight local
CompareChecking fit…
Alibaba · Hosted API
Price $0.12/M input · $0.46/M output
Context 256K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · qwen/qwen3-vl-8b-instruct
Moderate estimate $3.71/mo
Best for: Local visual perception, image/question tasks, and OCR-style workflows.
Not ideal for: Frontier multimodal reasoning.
value Vision Tools image-understanding multimodal local low-cost
CompareChecking fit…
Alibaba · Hosted API
Price $0.14/M input · $1.00/M output
Context 262K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · qwen/qwen3.6-35b-a3b
Moderate estimate $5.80/mo Architecture Gated DeltaNet and gated-attention hybrid with MoE layers
Best for: Coding help, local agents, multimodal notes, and tool-heavy workflows.
Not ideal for: Tiny/on-device setups.
value Vision Tools Reasoning text code image-understanding tool-use agentic open-weight local
CompareChecking fit…
Alibaba · Model download
Price Input not observed · Output not observed
Context 262K Availability Not observed Offer LM Studio
More technical evidence Exact ID · lmstudio-community/Qwen3.6-35B-A3B-GGUF:Q4_K_M
Moderate estimate — Artifact GGUF · Q4_K_M Reported download 21.2 GB Minimum system memory 20 GB · LM Studio Architecture Gated DeltaNet and gated-attention hybrid with MoE layers
Reported download size is storage evidence, not a RAM or VRAM requirement. The separate 20 GB · LM Studio minimum is source-backed, but still does not prove complete hardware fit.
Best for: Coding help, local agents, multimodal notes, and tool-heavy workflows.
Not ideal for: Tiny/on-device setups.
value Vision Tools Reasoning text code image-understanding tool-use agentic open-weight local
CompareChecking fit…
Alibaba · Model download
Price Input not observed · Output not observed
Context 262K Availability Not observed Offer LM Studio
More technical evidence Exact ID · lmstudio-community/Qwen3.6-35B-A3B-GGUF:Q6_K
Moderate estimate — Artifact GGUF · Q6_K Reported download 28.5 GB Minimum system memory 20 GB · LM Studio Architecture Gated DeltaNet and gated-attention hybrid with MoE layers
Reported download size is storage evidence, not a RAM or VRAM requirement. The separate 20 GB · LM Studio minimum is source-backed, but still does not prove complete hardware fit.
Best for: Coding help, local agents, multimodal notes, and tool-heavy workflows.
Not ideal for: Tiny/on-device setups.
value Vision Tools Reasoning text code image-understanding tool-use agentic open-weight local
CompareChecking fit…
Alibaba · Model download
Price Input not observed · Output not observed
Context 262K Availability Not observed Offer LM Studio
More technical evidence Exact ID · lmstudio-community/Qwen3.6-35B-A3B-GGUF:Q8_0
Moderate estimate — Artifact GGUF · Q8_0 Reported download 36.9 GB Minimum system memory 20 GB · LM Studio Architecture Gated DeltaNet and gated-attention hybrid with MoE layers
Reported download size is storage evidence, not a RAM or VRAM requirement. The separate 20 GB · LM Studio minimum is source-backed, but still does not prove complete hardware fit.
Best for: Coding help, local agents, multimodal notes, and tool-heavy workflows.
Not ideal for: Tiny/on-device setups.
value Vision Tools Reasoning text code image-understanding tool-use agentic open-weight local
CompareChecking fit…
Alibaba · Model download
Price Input not observed · Output not observed
Context 262K Availability Not observed Offer LM Studio
More technical evidence Exact ID · lmstudio-community/Qwen3.6-35B-A3B-MLX-4bit
Moderate estimate — Artifact MLX · 4-bit · Apple Silicon Reported download 20.4 GB Minimum system memory 20 GB · LM Studio Architecture Gated DeltaNet and gated-attention hybrid with MoE layers
Reviewed exact Runner support
LM Studio · 0.4.13 Build 1 · requires macOS + Apple Silicon
Exact reviewed pairs only; this does not imply support for adjacent artifacts or complete hardware fit.
Reported download size is storage evidence, not a RAM or VRAM requirement. The separate 20 GB · LM Studio minimum is source-backed, but still does not prove complete hardware fit.
Best for: Coding help, local agents, multimodal notes, and tool-heavy workflows.
Not ideal for: Tiny/on-device setups.
value Vision Tools Reasoning text code image-understanding tool-use agentic open-weight local
CompareChecking fit…
Alibaba · Model download
Price Input not observed · Output not observed
Context 262K Availability Not observed Offer LM Studio
More technical evidence Exact ID · lmstudio-community/Qwen3.6-35B-A3B-MLX-6bit
Moderate estimate — Artifact MLX · 6-bit · Apple Silicon Reported download 29.1 GB Minimum system memory 20 GB · LM Studio Architecture Gated DeltaNet and gated-attention hybrid with MoE layers
Reviewed exact Runner support
LM Studio · 0.4.13 Build 1 · requires macOS + Apple Silicon
Exact reviewed pairs only; this does not imply support for adjacent artifacts or complete hardware fit.
Reported download size is storage evidence, not a RAM or VRAM requirement. The separate 20 GB · LM Studio minimum is source-backed, but still does not prove complete hardware fit.
Best for: Coding help, local agents, multimodal notes, and tool-heavy workflows.
Not ideal for: Tiny/on-device setups.
value Vision Tools Reasoning text code image-understanding tool-use agentic open-weight local
CompareChecking fit…
Alibaba · Model download
Price Input not observed · Output not observed
Context 262K Availability Not observed Offer LM Studio
More technical evidence Exact ID · lmstudio-community/Qwen3.6-35B-A3B-MLX-8bit
Moderate estimate — Artifact MLX · 8-bit · Apple Silicon Reported download 37.7 GB Minimum system memory 20 GB · LM Studio Architecture Gated DeltaNet and gated-attention hybrid with MoE layers
Reviewed exact Runner support
LM Studio · 0.4.13 Build 1 · requires macOS + Apple Silicon
Exact reviewed pairs only; this does not imply support for adjacent artifacts or complete hardware fit.
Reported download size is storage evidence, not a RAM or VRAM requirement. The separate 20 GB · LM Studio minimum is source-backed, but still does not prove complete hardware fit.
Best for: Coding help, local agents, multimodal notes, and tool-heavy workflows.
Not ideal for: Tiny/on-device setups.
value Vision Tools Reasoning text code image-understanding tool-use agentic open-weight local
CompareChecking fit…
Amazon · Hosted API
Price $0.06/M input · $0.24/M output
Context 300K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · amazon/nova-lite-v1
Moderate estimate $1.92/mo
Best for: High-volume summaries, extraction, and AWS-native workloads.
Not ideal for: Complex reasoning or agentic coding.
budget Vision Tools text image-understanding summarization classification extraction low-cost
CompareChecking fit…
OpenAI · Hosted API
Price $0.05/M input · $0.40/M output
Context 400K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · openai/gpt-5-nano
Moderate estimate $2.20/mo
Best for: Routing, classification, short extraction, and simple Q&A.
Not ideal for: Reasoning-heavy or long-form tasks.
budget Vision Tools Reasoning text classification extraction summarization fast-response low-cost
CompareChecking fit…
OpenAI · Hosted API
Price $0.03/M input · $0.17/M output
Context 131K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · openai/gpt-oss-120b
Moderate estimate $1.11/mo Parameters 117B total · 5.1B active License Apache License 2.0
Best for: Self-hosting, privacy-sensitive workflows, and customization.
Not ideal for: Managed convenience or multimodal work.
budget Tools Reasoning text code reasoning open-weight self-hostable privacy-sensitive
CompareChecking fit…
OpenAI · Hosted API
Price $0.03/M input · $0.13/M output
Context 131K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · openai/gpt-oss-20b
Moderate estimate $0.99/mo Parameters 21B total · 3.6B active License Apache License 2.0
Best for: Private local assistant, reasoning with tool support, and low-cost hosted routes.
Not ideal for: Multimodal work.
budget Tools Reasoning text reasoning tool-use open-weight self-hostable privacy-sensitive
CompareChecking fit…
IBM · Hosted API
Price $0.05/M input · $0.10/M output
Context 131K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · ibm-granite/granite-4.1-8b
Moderate estimate $1.30/mo Parameters 8B total Architecture Decoder-only dense Transformer License Apache License 2.0
Best for: Local enterprise assistant, instruction following, and JSON/tool tasks.
Not ideal for: Hard reasoning or multimodal work.
budget Tools text tool-use extraction open-weight local
CompareChecking fit…
NVIDIA · Hosted API
Price $0.05/M input · $0.20/M output
Context 262K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · nvidia/nemotron-3-nano-30b-a3b
Moderate estimate $1.60/mo
Best for: Low-latency local reasoning, chat, and tool workflows.
Not ideal for: Multimodal or vision tasks.
budget Tools Reasoning text reasoning tool-use open-weight fast-response
CompareChecking fit…
Microsoft · Hosted API
Price $0.07/M input · $0.14/M output
Context 16K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · microsoft/phi-4
Moderate estimate $1.82/mo Parameters 14B total Architecture Dense decoder-only Transformer License MIT License
Best for: Small local reasoning, math, and short code help.
Not ideal for: Long-context or multimodal work.
budget text reasoning fast-response local open-weight
CompareChecking fit…
Microsoft · Model download
Price Input not observed · Output not observed
Context 16K Availability Not observed Offer LM Studio
More technical evidence Exact ID · lmstudio-community/phi-4-GGUF:Q4_K_M
Moderate estimate — Artifact GGUF · Q4_K_M Reported download 9.05 GB Minimum system memory 8 GB · LM Studio Parameters 14B total Architecture Dense decoder-only Transformer License MIT License
Reported download size is storage evidence, not a RAM or VRAM requirement. The separate 8 GB · LM Studio minimum is source-backed, but still does not prove complete hardware fit.
Best for: Small local reasoning, math, and short code help.
Not ideal for: Long-context or multimodal work.
budget text reasoning fast-response local open-weight
CompareChecking fit…
Alibaba · Hosted API
Price $0.07/M input · $0.28/M output
Context 160K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · qwen/qwen3-coder-30b-a3b-instruct
Moderate estimate $2.24/mo Architecture Causal language model with 128 experts and 8 activated experts
Best for: Local code help and software-agent experiments.
Not ideal for: Frontier coding benchmarks.
budget Tools code tool-use agentic open-weight local
CompareChecking fit…
Alibaba · Hosted API
Price $0.07/M input · $0.26/M output
Context 1M Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · qwen/qwen3.5-flash-02-23
Moderate estimate $2.08/mo
Best for: Bulk processing, extraction, classification, and simple coding.
Not ideal for: Deep reasoning or high-quality long-form writing.
budget Vision Tools Reasoning text code summarization classification extraction fast-response low-cost
CompareChecking fit…
DeepSeek · Hosted API
Price Input not observed · Output not observed
Context 1M Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · deepseek/deepseek-v4-flash
Moderate estimate — Maximum output 384,000 tokens
Published provider ceiling, not guaranteed visible answer length; reasoning, input, endpoint, or account conditions may reduce it.
Best for: Bulk summarization, extraction, routing, and classification.
Not ideal for: Complex reasoning or coding.
Tools Reasoning text summarization classification extraction fast-response low-cost
CompareChecking fit…
Mistral · Hosted API
Price $0.40/M input · $2.00/M output
Context 262K Availability Unavailable · stale 22d old Offer OpenRouter
More technical evidence Exact ID · mistralai/devstral-2512
Moderate estimate $14.00/mo
Best for: Codebase exploration, editing workflows, and agentic coding.
Not ideal for: General chat or lightweight summarization.
Tools code tool-use agentic planning open-weight
CompareChecking fit…
Google · Hosted API
Price $0.14/M input · $0.34/M output
Context 262K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · google/gemma-4-31b-it
Moderate estimate $3.82/mo
Best for: Local/private assistants, multimodal summaries, and document/image understanding.
Not ideal for: Lowest-resource laptops or top-tier coding agents.
Vision Tools Reasoning text image-understanding multimodal open-weight self-hostable local
CompareChecking fit…
Alibaba · Hosted API
Price $0.30/M input · $2.00/M output
Context 262K Availability Available · stale 22d old Offer OpenRouter
More technical evidence Exact ID · qwen/qwen3.6-27b
Moderate estimate $12.00/mo
Best for: Local coding, reasoning, research drafts, and multimodal input.
Not ideal for: Very low-memory setups.
Vision Tools Reasoning text code reasoning multimodal open-weight local
CompareChecking fit…