What the table is really for
Model choice is rarely just about one benchmark. The right answer depends on latency, data sensitivity, tool reliability, prompt complexity, expected volume, and how painful migration would be later.
A practical model explorer for teams choosing between hosted APIs, private deployments, and hybrid AI agent stacks. Sort the table, filter for fit, export a CSV, then bring the shortlist into a real architecture conversation.
Use this as a first-pass comparison. Pricing and model cards change, so confirm vendor details before purchasing infrastructure or committing production budget.
| Name | Company | Deployment | Input / 1M | Output / 1M | Context | Reasoning | Agent Fit | Best Use | Risk Note |
|---|
Model choice is rarely just about one benchmark. The right answer depends on latency, data sensitivity, tool reliability, prompt complexity, expected volume, and how painful migration would be later.
Use these lanes when you need a plain-English starting point.
Herb can review your workload, data sensitivity, latency needs, and budget, then recommend a model path that fits the actual system you are building.