Skip to content

Generation models

What murakumo.cloud can generate, and how well each model actually works. The fleet runs open-weight models on its own GPUs at $0/h and reaches hosted APIs for the jobs the fleet cannot do; both are listed here together.

These statuses are measured, not claimed

Every model id, status, price and timing on this page is read from the same EDN file the scheduler and the workers run on — there is no second table to fall out of date. A model that was measured broken stays listed and is labelled broken, because deleting it is how one stayed the default for eight days while returning empty video. Models with no recorded measurement say unverified rather than borrowing credibility from the ones that have one.

Host and distribute a model

murakumo.cloud now includes a public model repository API, separate from the fleet's runnable-model catalogue below. Publishers upload immutable content-addressed blocks, then publish a DAG-CBOR revision manifest with model card, license, file hashes and provenance. Public reads and byte-range downloads need no token; repository, block and revision writes require a DID-bound murakumo token with the model-hosting scope.

This is the hosting plane: durable weights and versioned metadata. The tables below are the execution plane: models currently wired into the Murakumo generation fleet. Hosting a revision does not claim that the fleet can execute it.

Model Hub

Publish a repository and immutable revision directly from the browser, import a pinned Hugging Face revision, request gated access, or contribute examples and reviews. Tokens remain in this page's memory and are never saved to browser storage.

Publish model

Import from Hugging Face

Request fleet promotion

Request gated access

Rate or comment

Publish example

Community ranking

  1. Loading…

The browser publisher uses 95 MiB CID blocks and intentionally caps one-file publishing at 512 MiB because Web Crypto does not expose a streaming SHA-256 API. Hugging Face imports run server-side in 16 MiB resumable chunks, pin the immutable upstream commit and LFS hashes, and remain quarantined until the static scan passes or is overridden.

Video

Video models
Model Runs on Status Modes Measured cost Speed
10eros-max self-hosted ($0/h fleet) verified t2v · i2v · ref2v · audio $0.00 (fleet GPU) 135s / clip
hokusai hosted (fal.ai) pending t2v · i2v · audio — —
hunyuanvideo-1.5 self-hosted ($0/h fleet) verified t2v · i2v $0.00 (fleet GPU) 850s / clip
ltx-2.3 self-hosted ($0/h fleet) broken t2v · audio $0.00 (fleet GPU) 140s / clip
minimax-h3 · default self-hosted ($0/h fleet) verified t2v · i2v · ref2v · audio $0.00 (fleet GPU) 6038s / clip
seedance-2.0 hosted (fal.ai) production t2v · i2v · ref2v — —
seedance-2.0-fast hosted (fal.ai) production t2v · i2v · ref2v $0.22 / page 90s / clip
veo-3.1 hosted (fal.ai) unverified t2v · i2v · audio — —
veo-3.1-fast hosted (fal.ai) unverified t2v · i2v — —
veo-3.1-silent hosted (fal.ai) unverified t2v · i2v — —
wan-2.7 hosted (fal.ai) unverified t2v · i2v · ref2v — —
wan-dancer-14b self-hosted ($0/h fleet) pending i2v — —
wan2.2-ti2v-5b self-hosted ($0/h fleet) verified t2v · i2v $0.00 (fleet GPU) 135s / clip
hunyuanvideo-1.5
Best OSS output quality here: monochrome-manga by default with no prompting for it. Also the slowest by a wide margin.
ltx-2.3
Was the default until 2026-08-03, when image-to-video was measured returning status=done with every frame black (luma mean 0, std 0). Text-to -video passed an end-to-end run on 2026-07-16 and is still selectable; image-to-video is refused rather than silently returning an empty clip.
seedance-2.0
Hosted reference-to-video, up to 4K, up to 9 reference images. Verified against the live API on 2026-08-03.
seedance-2.0-fast
The production path for comic work as of 2026-08-03. Multi-shot actually works: one 15-second clip came back containing the four requested shots in order — a whole page from a single generation, which is why per-page cost is a fair unit here.
wan-dancer-14b
Dance / cosplay image-to-video. Weights are still being fetched by the disk-gated watcher and the graph is not wired, so selecting it fails closed instead of falling back to something else.
wan2.2-ti2v-5b
Default. Pen-and-ink look down to the screentone grain; character consistency holds from a single reference. Output is a fixed 768x448 landscape, so tall panels get cropped — measured 2026-08-03.

Cost is per manuscript page rather than per clip because the two paths divide the work differently: the fleet models take one clip per shot, while Seedance's multi-shot returns a page's worth in a single generation. Per-clip pricing would flatter the hosted path by 4x.

Image

Image models
Model Runs on Status Notes
animagine-xl-4.0 · default self-hosted ($0/h fleet) verified Default image model. Anime/illustration checkpoint.
flux.1-dev self-hosted ($0/h fleet) unverified General-purpose diffusion. Flux.1, not Flux 2.
qwen-image self-hosted ($0/h fleet) unverified Qwen-Image. Strong at rendering text inside the image.
sdxl-turbo self-hosted ($0/h fleet) unverified Few-step SDXL for drafts and previews.

3D, rig and VRM

3D / rig / compose models
Model Runs on Status Notes
hunyuan3d-2.1 self-hosted ($0/h fleet) unverified Alternative image → 3D mesh.
trellis · default self-hosted ($0/h fleet) unverified Image → 3D mesh. Output glTF feeds kami-nerf / gsplat / terrain.
unirig self-hosted ($0/h fleet) unverified Auto-rig: mesh → humanoid skeleton + skin weights, exported as VRM 1.0.
bbox-humanoid-v1 self-hosted ($0/h fleet) unverified Explicit CPU fallback after a UniRig failure. Structurally valid VRM; deformation quality is degraded and is surfaced as rig:fallback.
kisekae-v1 self-hosted ($0/h fleet) unverified VRM compositor: reconnect skin, add/remove mesh, apply material and expression.

Audio, voice, motion and effects

Audio / voice / motion / effect models
Model Runs on Status Notes
ace-step self-hosted ($0/h fleet) production Default music model. ACE-Step-v1-3.5B on gad.
minimax-music3 hosted (fal.ai) pending —
musicgen-small self-hosted ($0/h fleet) verified MusicGen Small for music and sound effects.
cosyvoice2 self-hosted ($0/h fleet) unverified Menu only. Declared Japanese voice. No measured public TTS path; POST /api/v1/audio/speech returns 501 until a local backend is bound.
kokoro self-hosted ($0/h fleet) unverified Menu only. Declared lighter voice. No measured public TTS path; POST /api/v1/audio/speech returns 501 until a local backend is bound.
edn-motion-v1 self-hosted ($0/h fleet) unverified Validated humanoid motion clips as EDN on the Mac control worker.
bounded-preset-v1 self-hosted ($0/h fleet) unverified Deterministic allowlisted effects (sparkle / hearts / rain / smoke / impact). A prompt never becomes executable shader code.

Hugging Face performance snapshot

The current Murakumo text model and the comparison set published in its official Hugging Face model card. This is not a ranking of every recently updated Hub upload.

Snapshot 2026-08-29. Scores below are publisher-reported, not Murakumo measurements. Source: Qwen/Qwen3.8-27B at revision 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0. Values are publisher-reported. A row is comparable only within the evaluation setup and notes in the pinned model card.

Publisher-reported benchmark scores from the pinned model card
Area Benchmark Qwen3.8-27B Qwen3.6-27B Qwen3.7-Plus Muse Glimmer-30B Opus4.6 Max
Coding Terminal Bench 2.1 73 63.4 64 51.7 78.2
Coding SWE-bench Pro 61.7 53.5 57.6 51.2 53.4
General IFBench 79.5 69.1 79.1 77 62.5
General GPQA Diamond 89.2 87.8 90.3 83.5 91.3
General LiveCodeBench v6 90.3 83.9 89.6 — 88.8
Agentic multimodal OSWorld-Verified 84.3 63.9 73.3 65.9 72.7
Agentic multimodal WebArena-Verified 64.8 48.8 55.3 — —
Multimodal OmniDocBench 1.5 91.1 89.4 91.4 75.8 86.6

Murakumo route throughput snapshot

All currently advertised text routes use the Qwen3.8-27B family, so this table compares execution routes, not different model capabilities. The fixed request used 27 prompt tokens, up to 192 completion tokens, and temperature 0.

One fixed request against the live API — 2026-08-29T01:23:08Z
Requested route Resolved model Status Runs Prompt eval Generation Result
murakumo-main qwen3.8-27b-throughput-b70 measured 1 22.89 tok/s 38.62 tok/s n=1
qwen3.8-27b-fastmtp-aggressive murakumo-main route mismatch 1 11.4 tok/s 2.44 tok/s Response identified itself as murakumo-main.
qwen3.8-27b-throughput — unavailable 1 — — Fleet response was not valid JSON.
qwen3.8-27b-throughput-5090 — unavailable 1 — — Fleet response ended before a valid JSON document was returned.
qwen3.8-27b-throughput-b70 qwen3.8-27b-throughput-b70 measured 2 21.71–53.94 tok/s 38.75 tok/s n=2; generation range 38.73–38.76 tok/s.

Operational snapshot only (n=1 unless noted), not a controlled performance qualification. Generation tok/s is the backend-reported predicted_per_second, not end-to-end wall-clock throughput. A missing tok/s value means the route did not return a valid completion; it is not treated as zero throughput. tok/s depends on hardware, runtime, quantization, context, cache state, concurrency and sampling settings.

Inference optimization taxonomy

The text and omni models the inference fleet serves, classified by which optimization techniques their architecture carries and which serving-time techniques are applied to them. This is a different plane from the generation catalogue above: those models produce images, video, 3D and audio; these answer inference routes on api.murakumo.cloud.

Architecture facts are read from each publisher's Hugging Face config.json at the pinned revision; serving facts come from the murakumo model registry. Anything marked community-reported or our analysis has not been measured on this fleet. A technique chip in the Serving column is green only when the deployment actually applies it today; grey chips are carried capability or candidates, not claims.

Scope
Technique
Optimization techniques and where they live
Technique Class What it does
moe
MoE (Mixture of Experts)
architecture Only a routed subset of expert FFNs runs per token; total parameters grow without proportional per-token compute.
gated-deltanet
Gated DeltaNet (linear attention)
architecture Linear-attention token mixer with a gated delta-rule state update; constant-memory recurrent state replaces the KV cache on those layers.
kda
KDA (Kimi Delta Attention)
architecture Gated delta-rule linear attention variant (Kimi-style); GLM-5.3-Flash uses it on 34 of 45 layers.
mla
MLA (Multi-head Latent Attention)
architecture Low-rank latent KV projection (DeepSeek-style) shrinking the per-token KV cache on attention layers.
dsa
DSA (DeepSeek Sparse Attention)
architecture Full-attention layers attend only to a learned top-k subset of tokens chosen by a lightweight indexer.
qsa
QSA (Qwen Sparse Attention)
architecture Qwen's sparse attention: selects learned micro-blocks rather than individual tokens (community-reported from the Qwen3.8-Flash-Next release; no public paper yet).
swa-hybrid
Sliding-window / global hybrid
architecture Interleaved local sliding-window and full global attention layers; the final layer is global.
mhc
mHC (Manifold-Constrained Hyper-Connections)
architecture Four parallel residual streams with a manifold-constrained mixing matrix (GLM-5.3-Flash: config mhc=true, hc_mult=4).
indexpool
IndexPool
architecture Compresses the sparse-attention indexer's key vectors by weighted pooling (GLM-5.3-Flash: index_kpool_compress=true) to cut indexer latency at 1M context.
mtp
MTP (multi-token prediction)
architecture The checkpoint ships dedicated draft layer(s) that propose several tokens per step; the verifier keeps only the accepted ones.
flash-attn
FlashAttention
serving IO-aware exact attention kernel. llama.cpp: -fa on. Enabled on the xavier deployment.
spec-mtp
Speculative decoding: draft-mtp
speculative llama.cpp --spec-type draft-mtp: self-speculative decode using the checkpoint's own MTP layers. Needs an MTP-bearing GGUF (grafted builds exist for Qwen3.6).
spec-dflash
Speculative decoding: draft-dflash
speculative llama.cpp draft-dflash: a block-diffusion-style drafter paired with the target model.
spec-ngram
Speculative decoding: ngram family
speculative Draft-model-free lookup drafting: ngram-simple, ngram-map-k, ngram-map-k4v, ngram-mod, ngram-cache. --spec-default enables ngram-mod.
kv-quant
KV cache quantization
quantization q8_0 / f16 KV cache types trade context length against accuracy; community recommendation for GLM-5.3-Flash is q8_0.
gguf-quant
GGUF K-quant
quantization Q3_K_M / Q4_K_M / Q6_K and Unsloth Dynamic quants size the model to the accelerator.
fp8
FP8 checkpoint
quantization Native fp8 weight quantization shipped by the publisher (GLM-5.3-Flash quantization_config quant_method=fp8).
nvfp4-awq
NVFP4 / AWQ encoders
quantization 4-bit text-encoder quants used in the ComfyUI serving graph.
turbo-hybrid
TURBO-hybrid step distillation
diffusion Few-step diffusion distillation hybrid (the 10Eros-Max H3 finetune).
rpc-ring
Distributed RPC ring
distributed llama.cpp RPC distributes layers across a worker ring; the qwen-agentworld deployment serves with 9 Metal workers + head.
parallel-slots
Parallel slots
serving llama.cpp --parallel N: multiple concurrent sequences share one loaded model.
Served models, classified
Model Params Architecture facts Techniques (green = applied) Deployments
basho-320-18b
GLM-5.3-Flash (abliterated GGUF)
HF model card ↗
320B · 18B active 3 × linear (KDA) : 1 × sparse (DSA/MLA), 45 layers moe kda mla dsa mhc indexpool mtp fp8 gguf-quant parallel-slots spec-mtp kv-quant spec-dflash qsa Modal B200 single-GPU, ctx 32768
qwen3.6-35b-a3b
Qwen3.6 (gad standalone + xavier + RPC ring)
HF model card ↗
35B · 3B active 10 × (3 × linear : 1 × full), 40 layers moe gated-deltanet mtp spec-mtp flash-attn rpc-ring parallel-slots gguf-quant gad standalone GPU
xavier Jetson AGX (flash-attn on)
qwen-agentworld RPC ring ×9 workers
qwen3.8-27b (murakumo-main target)
Qwen3.8-27B throughput slot (b70)
HF model card ↗
27B config.json architecture fields not yet pulled for this card; pending parallel-slots spec-ngram spec-mtp qsa throughput-b70 slot, parallel=2, ctx 262144, vision
gemma-4-12b-it
Gemma 4 12B (retired from serving)
HF model card ↗
12B sliding-window / global hybrid, final layer global swa-hybrid —
hokusai (MiniMax H3 video)
MiniMax H3 video (10Eros-Max finetune)
HF model card ↗
20B Omni-modal video system: Context-IR → H3-Base → upscaler; separate text encoder (NVFP4/AWQ) and VAE (video fp16 / audio fp32). turbo-hybrid nvfp4-awq gad (gfx1151 ROCm) ~2200 s/clip
Modal H100 ~197 s/clip, weights volume 42.5 GB, cold start ~28 s

basho-320-18b — candidate optimizations (not yet applied)

  • spec-mtp — MTP + DFlash are reported effective on this architecture in community serving; the glm5next fork's spec path must be verified on our GPU before claiming it. community-reported — not measured here
  • kv-quant — q8_0 KV is the community recommendation; whether it fits B200 VRAM next to the 153 GB Q3_K_M weights must be measured. community-reported — not measured here
  • spec-dflash — Reported to combine with MTP in community benchmarks on hybrid-attention models. community-reported — not measured here
  • qsa — Not this checkpoint: QSA is reported for the Qwen3.8-Flash-Next line, not GLM. Listed so the filter can answer the question rather than leave it open. our analysis — not a measurement

qwen3.6-35b-a3b — candidate optimizations (not yet applied)

  • spec-mtp — MTP-grafted GGUFs (unsloth / am17an lineage) run --spec-type mtp; community reports draft_n=24 → accepted=19 on Strix Halo — the same gfx1151 class as gad. community-reported — not measured here

qwen3.8-27b (murakumo-main target) — candidate optimizations (not yet applied)

  • spec-ngram — Throughput-oriented slot: ngram-mod via --spec-default is the cheapest first probe (no draft weights needed). our analysis — not a measurement
  • spec-mtp — The Qwen3.8 line ships MTP layers; needs an MTP-bearing build on the slot before it can be claimed. our analysis — not a measurement
  • qsa — QSA is reported for the Qwen3.8-Flash-Next line; whether the b70 slot's checkpoint carries it is unverified. community-reported — not measured here

The vocabulary follows the publishers' names: MoE, Gated DeltaNet, KDA, MLA, DSA, QSA, sliding-window hybrids, mHC, IndexPool and MTP describe what the weights carry; FlashAttention, draft-mtp / draft-dflash / ngram speculative decoding, KV quantization, GGUF K-quants, fp8, RPC rings and parallel slots describe what a deployment chooses.

The measurement behind the video defaults

Six engines, one scenario, one session — the data the default model is chosen from, published rather than summarised.

Same scenario, six engines — gh-arc0-0-p17, 4 beats at 768x448, 2026-07-25
Engine Host Open weights $ / page s / clip Style fit Character Props Verdict
seedance-2.0-mini fal.ai no $0.22 90s 3 5 5 paid multishot champion
hunyuanvideo-1.5 gad yes $0.00 850s 5 5 3 oss quality champion
wan2.2-ti2v-5b gad yes $0.00 135s 5 4 2 oss balance
ltx-2.3 gad yes $0.00 140s 4 4 3 oss speed
animatediff-sdxl gad yes $0.00 140s 2 3 1 fragile
wan-2.7 — no — — — — — unavailable

Style fit / character consistency / prop intelligence are 1–5 scores from one measurement session on one scenario (monochrome manga + screentone, node: gad (AMD Ryzen AI Max+ 395 / gfx1151 / ROCm / ComfyUI 0.22.0)). They rank these six engines against each other on comic work; they are not general model benchmarks. Two ids in this table are not in the catalogue above: seedance-2.0-mini is the id this bench recorded for the fast hosted tier that the catalogue now serves as seedance-2.0-fast (same $0.22 per page, same 90s per clip), and animatediff-sdxl was evaluated and not adopted. Row names are left exactly as they were recorded.

Re-measured 2026-08-03

読切 P.12(4コマ) と P.21(4コマ) を genko の B4 原稿へ実際に組んだ — the run that moved the default.

  • ltx-2.3 — broken i2v ジョブが status=done を返すのに出力が全フレーム真っ黒(輝度 mean 0 / std 0、49frames 45KB)。 2026-07-25 の bench では :oss-speed だったが、2026-08-03 時点では成果物にならない。 無料経路を試すなら wan2.2 を使う。
  • wan2.2-ti2v-5b — works but unfit 動く。キャラ一致も良い(参照が単体なら)。ただし縦コマの原稿には合わない: ①出力が 768x448 の横長固定で、縦長のコマに入れると顔や主体が切れる ②i2v なので参照画像が1フレーム目になり、構図が寄り絵に固定される — 引きの絵や、参照と違う構図を作れない ③1クリップ1ショットなので 4コマ=4クリップ=約12分/頁 比率をコマに合わせて生成する改善は未検証(効く見込みはある)。
  • seedance-2.0-fast — production multi-shot が実際に効いた: duration 10 で 15秒クリップが返り、その中に 指示した4ショットが順に入っていた = 1クリップで原稿1頁ぶん。 1コマ1クリップで焼くと4倍払うことになる(実際に1頁 $0.88 払った)。

What is not here

Models people ask for that this fleet does not serve, and why.

Not available on murakumo.cloud
Model Why not
Seedream 5.0 · GPT Image 2 · Flux 2 · Nano Banana 2 / Pro · Ideogram 4 · Krea 2 · Qwen Image 2.0 Hosted image models. Reachable through the same fal.ai queue this fleet already uses for Seedance, but no adapter is written yet. Flux.1-dev and qwen-image (1.x) are the served predecessors of two of these — different checkpoints, not a version bump inside the existing graph.
Veo 3.1 · Kling 3.0 · PixVerse V6 · MiniMax H3 · Gemini Omni · Grok Imagine 1.5 · HappyHorse 1.0 · Happy Oyster Hosted video models, all present on fal.ai. What is missing is the input mapping: every one of them takes a `prompt` and returns `{video:{url}}`, but each spells duration and resolution its own way ("4s" vs "4" vs 4, 720p vs 768P), so each needs its own translation before it can be offered.
Wan 2.7 (on the $0/h fleet) Available hosted, NOT self-hostable. Official open weights stop at Wan 2.2, which is what the fleet serves — so 2.7 can never run on the free tier, only as a paid hosted call. Sites offering 2.7 weights for download are fakes.

Reading this catalogue from code

The same data, unstyled and always current: GET /api/v1/generation/catalog . No token required — it is a menu, not a capability. The document reports execute as menu-not-configured unless GET /api/v1 says generation-configured is true.

curl -s https://murakumo.cloud/api/v1/generation/catalog \
  | jq '{execute, speech, voice: .modalities.voice}'

Running a generation is gated and, on the public Worker observed 2026-08-27, not configured: POST /api/v1/generation needs a murakumo access token with the generation scope and a bound generation upstream. POST /api/v1/audio/speech is the OpenAI-shaped speech contract and returns 501 with configured: false until a local TTS backend can actually return audio. Voice ids (cosyvoice2, kokoro) stay on the menu as unverified. See the connection guide and the storefront.

Status vocabulary

What each status means
Status Meaning
production live-verified end to end, in production use
verified measured working on the fleet
broken measured NOT producing a usable artifact
pending weights or graph wiring not finished; selecting it fails closed
unverified declared and served, but no measurement is recorded here