# murakumo.cloud > Murakumo is a distributed inference and agent platform. This is the compact > machine-readable connection guide. The website and storefront live at > `https://murakumo.cloud`; the single public inference and agent API boundary > is `https://api.murakumo.cloud`. ## HTTPS is required Every endpoint below is HTTPS only. `http://` is answered with a redirect (301 for GET and HEAD, 308 for anything else so a POST keeps its method and body) and HSTS is set on the HTTPS response. Do not construct `http://` URLs for this host: the redirect is a safety net for a first visit, not a supported scheme. ## Public API boundary - Base URL for inference and agent clients: `https://api.murakumo.cloud` - Do not call `infer.murakumo.cloud`. It is an authenticated Cloudflare Tunnel origin used only by the API Worker to reach private fleet services. - Readiness: `GET https://api.murakumo.cloud/ready` - Model discovery: `GET https://api.murakumo.cloud/v1/models` - Agent capability discovery: `GET https://api.murakumo.cloud/v1/grok-bots` - Treat model registration and readiness as different facts. A model in `/v1/models` is a routable selector; clients should use bounded timeouts and must not describe it as live-verified until a real generation succeeds. ## OpenAI-compatible inference ### Chat Completions `POST https://api.murakumo.cloud/v1/chat/completions` The public first-value route accepts the OpenAI Chat Completions request shape and caps `max_tokens` at 2048. Always discover current model IDs first: the retired id `murakumo-main` now answers 410 (`model_retired`, retired 2026-09-20), so do not hard-code it. The free default is `murakumo/free` (OpenRouter free chain with the fleet's Mishima fallback, measured 200 on 2026-09-25); `mishima` is the fleet-hosted default measured 200 the same day. ```sh curl https://api.murakumo.cloud/v1/chat/completions \ -H 'content-type: application/json' \ -d '{ "model":"murakumo/free", "messages":[{"role":"user","content":"Reply exactly: ok"}], "max_tokens":64, "temperature":0 }' ``` ### RTX 5090 selector As observed from `/v1/models` on 2026-08-25, the registered selector is: - `qwen3.8-27b-throughput-5090` — advertised context window 32768 This selector is live-verified through the public API. On 2026-08-25 a scale-to-zero request returned HTTP 200 after 237.45 seconds, with response `.model` exactly `qwen3.8-27b-throughput-5090`; a warm request completed in 3.35 seconds. The gateway reserves up to 360 seconds for this RunPod cold-start path, so clients should allow at least 420 seconds end to end. Fall back explicitly to `mishima` if that bounded request fails. ```sh curl --max-time 420 https://api.murakumo.cloud/v1/chat/completions \ -H 'content-type: application/json' \ -d '{ "model":"qwen3.8-27b-throughput-5090", "messages":[{"role":"user","content":"Reply exactly: 5090-ok"}], "max_tokens":64, "temperature":0 }' ``` ## Tool calls For model-emitted function calls, use Chat Completions or the authenticated Anthropic Messages surface. Murakumo returns the requested tool call; the client executes the tool and sends the tool result in a follow-up request. Tool availability depends on the selected model and is not implied merely by its presence in `/v1/models`. ```sh curl --max-time 420 https://api.murakumo.cloud/v1/chat/completions \ -H 'content-type: application/json' \ -d '{ "model":"qwen3.8-27b-throughput-5090", "messages":[{"role":"user","content":"Use get_weather for Tokyo."}], "tools":[{ "type":"function", "function":{ "name":"get_weather", "description":"Get current weather", "parameters":{ "type":"object", "properties":{"location":{"type":"string"}}, "required":["location"] } } }], "tool_choice":{"type":"function","function":{"name":"get_weather"}}, "max_tokens":128, "temperature":0 }' ``` A successful model tool decision has `choices[0].finish_reason = "tool_calls"` and one or more entries under `choices[0].message.tool_calls`. ## Anthropic-compatible inference `POST https://api.murakumo.cloud/v1/messages` This surface translates Anthropic `tool_use` / `tool_result` blocks and streaming events to and from the fleet's OpenAI-compatible origins. It requires an authorized bearer or `x-api-key` accepted by Murakumo. ```sh export ANTHROPIC_BASE_URL=https://api.murakumo.cloud export ANTHROPIC_AUTH_TOKEN="$MURAKUMO_API_KEY" claude ``` Never put access tokens, operator signing secrets, or origin credentials in an agent configuration file committed to a repository. ## Responses and resident agents - Capability descriptor: `GET /v1/grok-bots` - Stateless Responses-compatible request: `POST /v1/responses` - Namespaced alias: `POST /v1/grok-bots/responses` - Namespaced Chat Completions alias: `POST /v1/grok-bots/chat/completions` - Resident runtime status: `GET /v1/grok-bots/runtime` - Resident bot management: `/v1/grok-bots/bots/*` The Responses-compatible transport is non-streaming and supports a portable subset of the OpenAI Responses request shape. The resident runtime is a bounded durable agent loop with persistent checkpoints and a capability allowlist. Management calls require a valid Murakumo service bearer. The live capability descriptor observed on 2026-08-25 advertised `murakumo-main` for the resident agent runtime. It did **not** advertise the 5090 selector. (The public id `murakumo-main` has since been retired, 2026-09-20; `/v1/responses` accepted `mishima` and `murakumo/free`, measured 2026-09-25.) Do not claim that the durable agent loop runs on RTX 5090 until `GET /v1/grok-bots` lists that model and an end-to-end agent tick succeeds. ```sh curl https://api.murakumo.cloud/v1/responses \ -H 'content-type: application/json' \ -d '{"model":"mishima","input":"Explain Murakumo in one sentence."}' ``` The canonical stateful agent service is `https://itonami.cloud/api/v1/grok-bots`; the `https://api.murakumo.cloud/v1/grok-bots` routes are compatibility aliases. Murakumo remains the inference provider. ## Authentication - `POST /v1/chat/completions` is the intentionally limited public first-value route. - `POST /v1/messages`, `POST /v1/embeddings`, slow jobs, and resident-agent management are authenticated surfaces. - Depending on the route, present either `Authorization: Bearer $MURAKUMO_API_KEY` or `x-api-key: $MURAKUMO_API_KEY`. - Tokens are capabilities. Use the shortest practical lifetime and scope; do not forward a client token to an origin or third-party tool. ## Image generation One image model is public and free on this same API boundary, with no account and no key. Published 2026-09-10. - Model id: `awai-network/hokusai` (the fleet's WAI checkpoint) - Route: `POST https://api.murakumo.cloud/v1/images/generations` - Listed by `GET https://api.murakumo.cloud/v1/models`, carrying its own `endpoint`, `kind` and `sizes`, so discovery is enough to call it. ```sh curl https://api.murakumo.cloud/v1/images/generations \ -H 'content-type: application/json' \ -d '{ "model":"awai-network/hokusai", "prompt":"a lighthouse on a rocky coast at dawn", "size":"768x768" }' ``` Returns `{"data":[{"b64_json":"..."}]}`. One image per request. `size` is an allowlist, not a range; an unsupported one is refused with the list. Measured 2026-09-10: about ten seconds for 768x768. It is refused on `/v1/chat/completions` with `400 wrong_endpoint_for_image_model`. That refusal is deliberate: an image model falling through a text route would reach the generic relay, which that boundary exists to prevent. ## Buying credits (top-up) When a free request is refused (error `reason` `viewer/insufficient-balance` → 402, `viewer/daily-cap-reached` or `viewer/balance-ceiling-reached` → 429 on the free lane, or a used-up daily allowance on `/go`), that is where a top-up begins. Where to send a human buyer, measured 2026-09-21: - Storefront and credits (top-up): `https://murakumo.cloud/portal/#store` — the working purchase page. It loads the account/storefront bundle (`/js/main.js`), sells prepaid credit SKUs, and checks out by card via Stripe; crypto is marked coming soon. Every site navigation links this same URL from its "クレジット" entry. - The bare fragment `https://murakumo.cloud/#store` is **not** the storefront: the home page carries no `#store` section and no checkout bundle (verified by response-body count on 2026-09-21). The storefront is served at `/portal/`. - `https://murakumo.cloud/pricing`, `/topup` and `/billing` are purchase aliases: since 2026-09-22 they answer 307 to `/portal/#store` (re-measured 2026-09-23; GET/HEAD only, bare spellings without a trailing slash). They are recognized redirects, not canonical URLs — cite `/portal/#store`. `console.murakumo.cloud` remains NXDOMAIN (re-measured 2026-09-23). `/portal/#store` is the only measured self-serve top-up surface. - Closing-loop reference (create did:key → Stripe checkout → mk1 token → API): `https://murakumo.cloud/docs/mvp-closed-loop.md` - Keyless paid alternative: x402 per-request settlement, section below. ## Credit units and reading a balance Murakumo sells and meters in **one unit: prepaid credits, 100 credits = $1 USD**. The rate is declared in code (`cloud-murakumo.x402-ledger/credits-per-usd` = 100, mirroring `local-murakumo.itonami/credits-per-usd`) and is the rate the storefront SKUs at `/portal/#store` are priced against. When a document says "credits" without qualification, it means this unit. Two other balance-looking numbers exist on other surfaces, and neither is this unit: - The ad-funded free lane reports a **micro-USD balance** (`balanceMicros`, 1,000,000 micros = $1; one free request costs 10000 micros = $0.01). It is scoped to the free lane, refilled by watching ads, and cannot be topped up with a card. - x402 settlement is **per-request USDC** ($0.01/request), never a stored balance. There is nothing to read back; each request is its own payment. To read a prepaid-credit balance: ``` GET https://murakumo.cloud/api/account/balance Authorization: Bearer ``` Returns `{"principal": ..., "credits": , "unit": {"name": "prepaid-credits", "creditsPerUsd": 100, "topUpUrl": "https://murakumo.cloud/portal/#store"}, "observedAt": ...}` with `cache-control: no-store`. An unauthenticated call answers 401; an unreachable ledger answers 503 (`balance_unavailable`). When `credits` reaches 0 the free public routes stop serving paid-shape requests and the next step is a top-up at `/portal/#store`. ## Other generation and the website Video, 3D, voice, and music generation remain a separate product surface documented by the website: - Catalogue: `GET https://murakumo.cloud/api/v1/generation/catalog` - Human-readable catalogue: `https://murakumo.cloud/models` - Execute generation: `POST https://murakumo.cloud/api/v1/generation` - Connection guide: `https://murakumo.cloud/docs` - Storefront and credits: `https://murakumo.cloud/portal/#store` The catalogue is a menu. Observed 2026-08-27: `GET https://murakumo.cloud/api/v1` reports `generation-configured: false` and `speech-configured: false`. `POST /api/v1/generation` exists as a gated execute path and is not a live backend until those flags are true. Do not treat catalog model ids as proof that execute works. Generation models are not returned by `api.murakumo.cloud/v1/models`; that endpoint is the inference model catalogue. ## Speech (OpenAI-shaped, not live) `POST https://murakumo.cloud/api/v1/audio/speech` This is the documented OpenAI-shaped TTS contract on the website Worker. It is not a live speech backend. Until `GET /api/v1` reports `speech-configured: true`, the route returns HTTP 501 with `configured: false`. There is no `/audio/speech` on `api.murakumo.cloud`. Catalog voice ids (`cosyvoice2`, `kokoro`) are unverified menu entries. Do not invent model IDs. `/go` hides Read aloud unless speech is configured. ```sh curl -i -X POST https://murakumo.cloud/api/v1/audio/speech \ -H 'content-type: application/json' \ -d '{"input":"hello"}' ``` A 501 body with `configured: false` is success for this contract. Do not retry it as if a fleet TTS worker were bound. ## Paying with x402 — and trying it for free on a testnet Discovery: `GET https://murakumo.cloud/.well-known/x402` lists every priced resource on both rails. Unpaid POSTs return HTTP 402. The live 402 paths are: - `POST https://murakumo.cloud/x402/v1/chat/completions` - `POST https://murakumo.cloud/x402/v1/infer-memory` - `POST https://murakumo.cloud/x402/v2/chat/completions` - `POST https://murakumo.cloud/x402/v2/infer-memory` Every priced resource here is sold twice: once on Base mainnet and once on **Base Sepolia**, at the same price. A 402 `accepts` list carries both rails — v1 uses `base` / `base-sepolia`; v2 uses CAIP-2 `eip155:8453` / `eip155:84532` with `extra.networkId` 8453 or 84532 matching that rail. Pick by your own network policy. The two differ in exactly two fields, and both matter: base asset 0x833589fC… payTo 0xA00366…D4E (real USDC) base-sepolia asset 0x036CbD53… payTo 0xD030410B…628A (test USDC) Send the wrong one and it will not settle: the network is part of the offer, not a detail of it. The `extra` EIP-712 domain differs too (`USD Coin` on mainnet, `USDC` on the testnet), which matters if you sign rather than transfer. **No testnet USDC?** Draw from `https://x402.nexus/faucet`. It hands out Base Sepolia gas and USDC, reports its pool as measured (how many draws are left, and which half is the binding constraint), and refuses honestly rather than pouring nothing quietly. The gate is a Biscuit from `auth.itonami.cloud`, rooted in a passkey sign-in. Nothing on the testnet rail costs anyone money: the payments settle to the faucet's own pool, so what it hands out returns when it is spent. ## Operational rules for agents 1. Use only `api.murakumo.cloud` for inference and agent calls. 2. Discover models and capabilities instead of hard-coding undocumented IDs. 3. Check `/ready`, then perform a small real request before claiming a model is available. 4. Assert the returned `.model`; HTTP 200 alone does not prove the requested backend served the request. 5. Use bounded timeouts and explicit fallback. Do not silently relabel fallback output as RTX 5090 output. 6. Never call `infer.murakumo.cloud` or expose its origin credentials. ## Advertising via MCP Endpoint: `https://murakumo.cloud/mcp` (Streamable HTTP, JSON responses). Connect an MCP client directly to this URL. No login or bearer token is required for public intake. Protocols: 2025-06-18 and 2025-03-26. GET returns 405 because no server-initiated SSE stream is offered. Tools: - `murakumo_ads_get_spec`: read placement, limits and review policy. - `murakumo_ads_validate`: validate `{campaign: {...}}` without saving. - `murakumo_ads_submit`: save an authorized application for automated Bot review. - `murakumo_ads_status`: read a private receipt; this does not verify delivery. Campaign required fields: `org`, `email`, `url` (HTTPS, no credentials), `creative` (1–120 characters). Optional string fields: `budget` (include currency), `start_date` (YYYY-MM-DD), `duration`, `notes`. The placement is `go-sidebar`, visible on the conversation screen on desktop and mobile. Do not include conversation text, prompts, generated output, secrets, tracking pixels or third-party ad scripts. Workflow: 1. Call get_spec and validate the campaign. 2. Generate 32 cryptographically random bytes as 64 lowercase hex characters (for example Python `secrets.token_hex(32)`). Save the key privately. 3. With advertiser authorization, call submit with `campaign`, `submission_key`, and `confirm_submission: true`. This acknowledges use of the submitted contact information for handling the application. 4. Persist the returned `application_id` and the original campaign and key. A retry with the identical normalized campaign and key converges on the same application. Changed campaign content creates a NEW application. 5. Call status with `application_id` and `submission_key`. Wait at least 60 seconds between polls; storage is eventually consistent. Keep the key out of URLs, public logs and shared transcripts. It authorizes receipt access only, and does not authenticate an advertiser's identity. The response includes machine-readable `structuredContent` and matching JSON text content. Handle `isError: true` even when HTTP status is 200. `receipt_unavailable` means missing receipt OR mismatched private key; a recent write may not yet be visible. Retry transient failures with unchanged input, never with a newly generated key. Concurrent retries converge on the same record but are not transactional; receipt timestamps may converge later. Intake is free. Pricing is an individual quote. Submission does not approve, activate, schedule or pay for an ad. There is no automated billing, guaranteed inventory, or impression report in these tools. Human form: https://murakumo.cloud/advertise ## Public Bot review of advertising All intake is scanned by the resident review Bot every five minutes, up to 2 applications per run. Queue size and KV propagation affect latency. Public process and event history: https://murakumo.cloud/advertise/review Machine-readable policy (including the exact model system prompt): https://murakumo.cloud/advertise/review/policy.json Paginated public events: https://murakumo.cloud/advertise/review/events.json Follow the returned cursor as the `cursor` query parameter until null. Bot execution status: https://murakumo.cloud/advertise/review/health.json Additional MCP tools: - `murakumo_ads_review_policy`: get full criteria, prompt and outcome codes. - `murakumo_ads_request_review`: with application_id and submission_key, request another Bot review, at most once per 24 hours per receipt. The existing status tool returns `review` (queued, approved, rejected, needs_revision, retry_pending, retry_exhausted), linked to the public case. The Bot checks creative bounds, public HTTPS/DNS and the landing-page response, then uses `qwen3.8-27b-throughput-b70` via the Murakumo public API to assess a bounded page excerpt. The returned model id and strict decision schema must match. Website content cannot supply review instructions or tool calls. Failures never become approval: transient errors retry at most three times; unverified model responses remain unverified. Modified content is submitted as a new application. A same-content appeal uses request_review and retains previous public events. New reviews can differ because pages and models change. Only fixed result codes, policy version, timestamps, Bot/model identifiers and evidence hashes are published. Contact email, original intake id, receipt capability, creative, URL, notes and raw model text are not published. The model sees only advertiser name, creative and an excerpt from the public page. The review is not a whole-site audit or a legal/identity certification. An approved review does not quote a price, authorize payment or start delivery.