API
Why PAYG if you can rent your own box
A self-hosted GLM on RunPod keeps the GPU billed even while you are idle — around $40 an hour. Here you pay for tokens. Packs start at $1 in Telegram, not an hourly box rental.
This is not a “we are the cheapest” page. It is billing shape: idle hours versus a million tokens. How a request is counted: what a request costs. Key path: how to start. Pick a head: which model. Six model cards live on pricing.
What you pay on your own box
While the instance is up, the clock runs. No prompts still heat the GPU. Drivers, disk, weight updates are yours. For occasional sessions that is more expensive than PAYG.
What you pay here
Input, output, cache per 1M tokens. GLM: $0.39 / $1.58 / $0.16. Minimum $0.005 per request. Empty model = GLM. Other rates — on /en/pricing/, do not average them.
Top up with $1, $5, $10, $25, $50, $100 packs in @uncensobot. No cashier on the site.
When your own box still wins
If you keep 24/7 load and want the hardware, rent can add up. Open a chat in the evening and quit — PAYG. Compare your idle hours to tokens, not someone else’s screenshot.
FAQ
Is there a free API?
No. The free slot is bot chat only: 1 per month, 2 with Premium, no search, no think.
Is $1 unlimited?
No. It is a balance pack. Tokens debit until it runs out.
Why is GLM cheaper than Large v2?
Different models, different per-1M rates. Read /en/pricing/.