Connecting to Murakumo
Use https://api.murakumo.cloud for every inference and agent call. murakumo.cloud is the website and storefront; infer.murakumo.cloud is an authenticated private origin and is not a client endpoint. The machine-readable guide is /llms.txt.
Public API
| Method | URL | Purpose |
|---|---|---|
| GET |
https://api.murakumo.cloud/ready
|
Gateway, fleet, and request-capacity readiness |
| GET |
https://api.murakumo.cloud/v1/models
|
Current inference selectors and context windows |
| POST |
https://api.murakumo.cloud/v1/chat/completions
|
OpenAI-compatible Chat Completions |
| POST |
https://api.murakumo.cloud/v1/messages
|
Authenticated Anthropic-compatible Messages |
| POST |
https://api.murakumo.cloud/v1/embeddings
|
Authenticated OpenAI-compatible embeddings |
| POST |
https://api.murakumo.cloud/v1/responses
|
Non-streaming portable Responses subset |
| GET |
https://api.murakumo.cloud/v1/grok-bots
|
Agent capability descriptor |
Web3 rails
Murakumo exposes the payment, identity, and content-addressing rails directly. Read the live offer before paying; a missing or empty offer is not permission to guess a recipient or asset.
| Method | URL | Purpose |
|---|---|---|
| GET |
https://murakumo.cloud/.well-known/x402
|
Machine-readable x402 resources, USDC amount, Base network, and pay-to address |
| POST |
https://murakumo.cloud/x402/v1/chat/completions
|
OpenAI-shaped inference; an unpaid request returns the HTTP 402 payment challenge |
| GET |
https://murakumo.cloud/api/v1/model-hosting
|
DID-published, CIDv1/DAG-CBOR model-hosting capability and limits |
| GET / HEAD |
https://murakumo.cloud/ipfs/{cid}
|
Read immutable content by CID |
| GET / HEAD |
https://murakumo.cloud/ipns/{name}
|
Resolve an IPNS name, then read its current CID |
curl -i -X POST https://murakumo.cloud/x402/v1/chat/completions \
-H 'content-type: application/json' \
-d '{"model":"murakumo-main","messages":[{"role":"user","content":"Reply exactly: ok"}],"max_tokens":16}'
Without X-PAYMENT, this request should stop at 402 and return the exact payment requirements. An x402 client settles the listed USDC requirement on Base and repeats the request with its payment proof. Do not treat a listed offer or a successful 402 challenge as proof that a paid inference completed; verify the final model response.
Model-hosting writes require a Murakumo token bound to did:key or did:web and scoped to model-hosting. Public repository and block reads require no token. The browser publisher and protocol are at /models/.
Chat Completions
The stable model alias is murakumo-main. The public first-value Chat Completions route caps max_tokens at 2048.
curl https://api.murakumo.cloud/v1/chat/completions \
-H 'content-type: application/json' \
-d '{"model":"murakumo-main","messages":[{"role":"user","content":"Reply exactly: ok"}],"max_tokens":64,"temperature":0}'
RTX 5090
The model catalogue observed on 2026-08-25 listed qwen3.8-27b-throughput-5090 with a 32768-token context window. It is live-verified through this public API: a scale-to-zero request returned HTTP 200 after 237.45 seconds with the exact requested model id, and a warm request completed in 3.35 seconds. The gateway allows up to 360 seconds for this cold-start path, so clients should use at least a 420-second end-to-end timeout and fall back explicitly to murakumo-main.
curl --max-time 420 https://api.murakumo.cloud/v1/chat/completions \
-H 'content-type: application/json' \
-d '{"model":"qwen3.8-27b-throughput-5090","messages":[{"role":"user","content":"Reply exactly: 5090-ok"}],"max_tokens":64,"temperature":0}'
Tool calls
Use Chat Completions for OpenAI-style tool_calls or authenticated Messages for Anthropic tool_use / tool_result. Murakumo returns the tool decision; your agent executes the tool and sends the result in a follow-up request. Tool support depends on the selected model and is not implied by model registration.
curl --max-time 420 https://api.murakumo.cloud/v1/chat/completions \
-H 'content-type: application/json' \
-d '{"model":"qwen3.8-27b-throughput-5090","messages":[{"role":"user","content":"Use get_weather for Tokyo."}],"tools":[{"type":"function","function":{"name":"get_weather","parameters":{"type":"object","properties":{"location":{"type":"string"}}},"required":["location"]}}],"tool_choice":{"type":"function","function":{"name":"get_weather"}},"max_tokens":128}'
Anthropic-compatible clients
The Messages surface requires an authorized Murakumo bearer or x-api-key. It translates streaming, thinking, and tool-use blocks to and from the fleet's OpenAI-compatible origins.
export ANTHROPIC_BASE_URL=https://api.murakumo.cloud export ANTHROPIC_AUTH_TOKEN=$MURAKUMO_API_KEY claude
Responses and resident agents
POST /v1/responses and POST /v1/grok-bots/responses provide a non-streaming portable Responses subset. The canonical stateful runtime is https://itonami.cloud/api/v1/grok-bots; the api.murakumo.cloud routes are compatibility aliases and Murakumo supplies inference. The live descriptor observed on 2026-08-25 advertises
murakumo-main for resident agents, not the 5090 selector. Do not claim a 5090 durable agent run until the descriptor lists it and an end-to-end agent tick succeeds. Resident-bot management requires a service bearer.
Authentication and safety
Chat Completions is an intentionally limited public first-value route. Messages, embeddings, slow jobs, and resident-agent management are authenticated. Depending on the route, send Authorization: Bearer *** or x-api-key: $MURAK...EY. Use short-lived, narrowly scoped capabilities and never commit tokens or origin credentials.
Image generation
One image model is public and free on the same API boundary, with no account and no key: awai-network/hokusai, the fleet's WAI checkpoint. It answers the OpenAI images shape and IS listed by api.murakumo.cloud/v1/models with the route it serves.
curl https://api.murakumo.cloud/v1/images/generations \
-H 'content-type: application/json' \
-d '{"model":"awai-network/hokusai",
"prompt":"a lighthouse on a rocky coast at dawn",
"size":"768x768"}'
Returns {"data":[{"b64_json":"..."}]}. One image per request; sizes are an allowlist, not a range. Measured 2026-09-10, a 768x768 render takes about ten seconds. It is refused on /v1/chat/completions with a 400 naming this route: an image model is not a chat model, and falling through would hand it to a relay.
Other generation, and the storefront
Video, 3D, voice, and music generation remain a separate website surface. See /models and the storefront. Those models are not listed by api.murakumo.cloud/v1/models.
GET /api/v1/generation/catalog is a menu. Observed 2026-08-27: GET /api/v1 reports generation-configured: false and speech-configured: false . Opening the catalog does not open execute.
Speech
The OpenAI-shaped speech contract is POST https://murakumo.cloud/api/v1/audio/speech. It is documented so clients can probe a stable path. It is not a live TTS backend. Until GET /api/v1 reports speech-configured: true, the route returns HTTP 501 with configured: false. Do not invent model ids. Voice ids in the catalog (cosyvoice2, kokoro) are unverified menu entries. There is no /voice page, no clone UI, and no avatar or lip-sync factory.
curl -i -X POST https://murakumo.cloud/api/v1/audio/speech \
-H 'content-type: application/json' \
-d '{"input":"hello"}'
A 501 body with configured: false is the honest answer. A 401 from this path would have meant the generic inference gate swallowed an unknown route — that is the failure this contract exists to stop. /go hides Read aloud unless speech is configured.