Models
Merki hosts a catalog of open and licensed models, and lets you bring your own key for providers you already use. Both routes are served by the same proprietary Merki inference stack, which is why throughput is high and prices are low.
Hosted models
Any model in the catalog can be called directly. You pay in credits at the listed per-million-token rates. Calls go through endpoints that are compatible with OpenAI and Anthropic clients. See Endpoints and compatibility.
Bring your own key
You can also route through your own provider key. See Bring your own key.
How to read the catalog
The catalog is a table. One row per model and quantization. Columns:
| Column | Meaning |
|---|---|
| Model | The model name, as published by its author. |
| Quant | The quantization Merki serves, for example FP4, FP8, BF16, INT4, GGUF. |
| Context | Maximum context window, in tokens. |
| Input $/M | Price per million input tokens, in US dollars. |
| Output $/M | Price per million output tokens, in US dollars. |
| Modalities | What the model accepts, for example text, image, video. |
| Tok/s | Output tokens per second, measured by Merki. |
| TTFT | Time to first token. |
| Uptime | Measured availability over the reporting period. |
| Model card | Links to the upstream model card and an independent speed benchmark. |
Prices reflect the proprietary Merki serving stack: every hosted price sits 25–35% below the cheapest published list price for the same model. Cache hits are billed at 50% of the listed input rate. See Pricing and Caching. Where a cell is not available it reads n/a.
Snapshots
The canonical table is the model catalog. Dated snapshots are frozen monthly, for example the August 2025 snapshot. The current page always holds the latest lineup.