Introduction
Harogo is a multi-model subscription service built for coding agents. The discounted subscription price is $4.9/month, with 100 discounted subscription slots currently available. Remaining slots and the actual billed price are shown on the subscription page. Once subscribed, you can use multiple coding models through a single Harogo account and API key. All models share the same subscription quota, so there’s no need to sign up for, fund, or manage separate model platforms.Why Harogo
Different models excel at different coding tasks. Using multiple model providers directly usually means juggling several accounts, API keys, balances, rate limits, and bills at once. Harogo brings together a range of models well-suited to coding agents into a single service. You’re free to switch between models based on the task, with everything billed against one shared subscription quota.Supported Models
Currently supported models:
The names and model IDs above are the routing identifiers Harogo uses internally. The model lineup, versions, and availability may change — check the console for the current list.
Usage and Limits
Harogo enforces the following usage limits:- Daily limit: $12
- Weekly limit: $30
- Monthly limit: $60
Estimated Request Volume
The table below estimates request counts per model based on typical Harogo usage patterns:
Each row assumes the quota for that time window is spent mostly on that one model. When you mix models, they draw from the same shared quota together, so the request counts in the table can’t simply be added up across rows.
Actual request volume is affected by factors like context length, cache hit rate, reasoning effort, and output length. A coding agent completing a single task also typically fires off multiple model requests.
Estimation Method
These estimates are based on the following typical request profile:- GLM-5.2 — about 700 input tokens, 52,000 cache-read tokens, and 150 output tokens per request.
- Qwen3.8 Max — about 8,731 input tokens, 31,831 cache-read tokens, and 2,029 output tokens per request.
- Kimi K3 — about 1,050 input tokens, 76,500 cache-read tokens, and 300 output tokens per request.
- MiniMax M3 — about 4,381 input tokens, 28,768 cache-read tokens, and 1,108 output tokens per request.
- DeepSeek V4 Pro — about 750 input tokens, 82,000 cache-read tokens, and 290 output tokens per request.
- DeepSeek V4 Flash — about 790 input tokens, 68,000 cache-read tokens, and 280 output tokens per request.
- Claude Opus 5 — about 580 input tokens, 386,230 cache-read tokens, 19,476 cache-write tokens, and 2,025 output tokens per request.
- Claude Sonnet 5 — about 159 input tokens, 129,804 cache-read tokens, 6,260 cache-write tokens, and 850 output tokens per request.
- Claude Fable 5 — about 399 input tokens, 315,393 cache-read tokens, 12,907 cache-write tokens, and 1,016 output tokens per request.
- Claude Haiku 4.5 — about 700 input tokens, 52,000 cache-read tokens, and 150 output tokens per request.
- GPT-5.6 Luna — about 45,726 input tokens, 10,720 cache-read tokens, and 1,905 output tokens per request.
- GPT-5.6 Sol — about 6,032 input tokens, 112,138 cache-read tokens, and 678 output tokens per request.
- GPT-5.6 Terra — about 6,825 input tokens, 69,476 cache-read tokens, and 734 output tokens per request.
Model Pricing and Usage Caps
The table below lists the reference API price per million tokens for each model, along with the per-model usage cap Harogo sets:
”—” means that item isn’t billed separately, or doesn’t apply to that model.
Model pricing, usage caps, and estimated request volumes may change as the service evolves — check the Harogo console and this page for the latest information.