ERNIE-Image × Civitai Orchestration API Complete Guide: Zero-GPU Cloud Text-to-Image with Buzz Pricing

Jul 2, 2026

ERNIE-Image × Civitai Orchestration API Complete Guide: Zero-GPU Cloud Text-to-Image with Buzz Pricing

Since ERNIE-Image went open source, developers have been searching for the simplest way to call this 8B DiT model. Local deployment requires a GPU, cloud deployment requires environment setup — until the arrival of Civitai Orchestration API, everything changed.

Civitai has integrated ERNIE-Image into its Orchestration API, allowing developers to directly call ERNIE-Image for image generation through a single REST API endpoint, with no local hardware, model downloads, or environment configuration needed. Even more exciting, Turbo mode boosts generation speed by 3-4x while reducing costs to one-third.

This article dives deep into the Civitai Orchestration API's ERNIE-Image integration — from API calls to cost optimization — helping you get started with cloud-based text-to-image generation quickly.

What Is Civitai Orchestration API

Civitai Orchestration API is a unified AI workflow orchestration platform supporting image, video, audio, and text generation. Its core philosophy is "workflows, not endpoints" — you describe the work you want done, the orchestrator picks a provider, routes the job, and streams results back. You don't manage any capacity.

Key features:

  • Multi-provider by default: FAL, Google, Bytedance, Civitai workers — the orchestrator races providers and selects the best fit for each job
  • Typed recipe catalog: One recipe per job type (video-gen, image-gen, upscaling, transcription, TTS…) with validated inputs and predictable outputs
  • Sync or async: Poll, subscribe, or wait inline with the wait= parameter. Webhooks supported for production integrations
  • MCP-native: Connect Claude Desktop, claude.ai, or any MCP-aware client to the same orchestrator

Two ERNIE-Image Modes in Civitai API

Civitai Orchestration API offers two modes for ERNIE-Image:

Mode Default Steps Default CFG Scale Best For
ernie (standard) 20 4 Default high-quality output, standard sampling
turbo 8 1 Distilled for speed — 3-4× faster, ~⅓ the Buzz cost

Critical tuning note: In Turbo mode, keep cfgScale at 1. Turbo is a distilled model and doesn't respond to classifier-free guidance like the standard variant. Raising the CFG Scale typically over-saturates/burns the output.

Complete API Call Guide

Basic Request Structure

Endpoint: POST https://orchestration.civitai.com/v2/consumer/workflows?wait=60

Headers:

Authorization: Bearer <TOKEN>
Content-Type: application/json

Standard Mode Example

{
  "steps": [
    {
      "$type": "imageGen",
      "input": {
        "engine": "comfy",
        "ecosystem": "ernie",
        "model": "ernie",
        "operation": "createImage",
        "prompt": "A red panda wearing a yellow rain jacket, cinematic soft light, highly detailed",
        "width": 1024,
        "height": 1024,
        "steps": 20,
        "cfgScale": 4,
        "sampler": "euler",
        "scheduler": "simple",
        "quantity": 1
      }
    }
  ]
}

Turbo Mode Example

Change "model" to "turbo", "steps" to 8, and "cfgScale" to 1:

{
  "steps": [
    {
      "$type": "imageGen",
      "input": {
        "engine": "comfy",
        "ecosystem": "ernie",
        "model": "turbo",
        "operation": "createImage",
        "prompt": "A red panda wearing a yellow rain jacket, cinematic soft light, highly detailed",
        "width": 1024,
        "height": 1024,
        "steps": 8,
        "cfgScale": 1,
        "sampler": "euler",
        "scheduler": "simple",
        "quantity": 1
      }
    }
  ]
}

Core Parameters Explained

Field Default Range Notes
prompt — ≤ 10,000 chars Natural language works best; handles complex scenes well
negativePrompt (none) ≤ 10,000 chars Optional. Shorter is better; defaults are already clean
width / height 1024 64–2048 (divisible by 16) Trained around 1024². Well-behaved near ~1 megapixel
steps 20 1–150 Diminishing returns past ~25. Turbo: keep 6–12
cfgScale 4 0–30 Sweet spot: 3–5. Turbo: lock at 1
sampler euler enum Model was tuned against euler
scheduler simple enum Standard scheduler
loras {} { airUrn: strength } Stack multiple. Only urn:air:ernie:lora:... work
quantity 1 1–12 Images per call
seed random int64 Pin for reproducibility

Buzz Pricing Model

Civitai uses the Buzz credit system for pay-per-use billing.

Pricing Formulas

Standard: total = 20 × (width × height / 1024²) × (steps / 20) × quantity

Turbo: total = 8 × (width × height / 1024²) × (steps / 8) × quantity

Real-World Cost Reference

Configuration Standard (Buzz) Turbo (Buzz)
1024² / default steps / quantity: 1 20 8
832×1216 / default steps / quantity: 1 ~20 ~8
1024² / default steps / quantity: 4 ~80 ~32
1024² / steps: 40(std) / steps: 16(turbo) ~40 ~16

Standard pricing is ~2.5× turbo at defaults. Recommendation: Use Turbo for prompt iteration, then switch to standard mode for final output.

Runtime Expectations

Variant Shape Expected Duration
ernie (standard) 1024² / 20 steps ~29 seconds
ernie (standard) 832×1216 / 20 steps ~27 seconds
turbo 1024² / 8 steps ~13 seconds

Response Format and Signed URLs

The API returns standard imageGen output with an images[] array:

{
  "status": "succeeded",
  "steps": [{
    "$type": "imageGen",
    "name": "$0",
    "status": "succeeded",
    "output": {
      "images": [{
        "id": "aa6e7228-68cd-4d15-b4d7-5005b2bfbac6-0.jpg",
        "width": 1024,
        "height": 1024,
        "url": "https://orchestration.civitai.com/v2/consumer/blobs/…?sig=…",
        "urlExpiresAt": "2027-04-15T17:18:54.3195353Z",
        "previewUrl": "https://orchestration.civitai.com/v2/consumer/blobs/…?sig=…",
        "available": true,
        "nsfwLevel": "pg13"
      }],
      "errors": []
    }
  }]
}

Important: url and previewUrl are signed and expire. For long-term use, re-fetch the workflow or call GetBlob for fresh URLs. nsfwLevel contains the moderation classification.

Async Calls and Polling Strategy

wait=60 covers single-image calls. For quantity > 1, large dimensions, or high steps, compute + queue will exceed 60 seconds.

Recommended polling strategy:

  1. Submit with wait=60
  2. If timed out, loop GET /v2/consumer/workflows/{id}?wait=60 until terminal status
  3. For production, register a webhook callback

LoRA Support

Civitai API only supports ERNIE-tagged LoRAs (urn:air:ernie:lora:...).

Add LoRAs to your request:

{
  "loras": {
    "urn:air:ernie:lora:civitai:12345@67890": 0.8
  }
}

Multiple LoRAs can be stacked by referencing different AIR URNs.

Use Cases

1. Prompt Iteration Testing

Use Turbo mode to quickly test different prompt combinations at minimal cost (8 Buzz/image), with results in just 13 seconds.

2. Batch Generation

Set quantity: 12 to generate up to 12 images per call — perfect for e-commerce product photography and social media content batch production.

3. API Integration

Integrate Civitai API into your own applications without managing GPU infrastructure. Webhook support makes it suitable for async production workflows.

4. Multi-Resolution Testing

ERNIE-Image performs best around 1024² but supports 64–2048 (divisible by 16). Use the API to quickly test different resolution effects.

FAQ

Q: What are the advantages over local deployment?

No GPU hardware needed, no model downloads, no environment configuration. API call equals image generation — perfect for rapid prototyping and small-scale production.

Q: How much quality loss in Turbo mode?

Turbo is optimized through DMD and RL distillation, achieving acceptable aesthetic quality in just 8 steps. For prompt iteration and draft generation, quality loss is negligible. Use standard mode for final output.

Q: How to get a Civitai API Token?

Register on the Civitai Developer platform and generate an API Token. See the Civitai Developer documentation for the detailed process.

Q: Does it support img2img or editing operations?

Currently only createImage (text-to-image) is supported. img2img, variant generation, and image editing are not yet available.

Summary

Civitai Orchestration API provides a zero-barrier cloud calling method for ERNIE-Image. Key advantages:

  • Zero hardware requirement: No GPU, no model download
  • Turbo acceleration: 3-4× speed, 1/3 cost
  • Batch capability: Up to 12 images per call
  • LoRA compatible: Supports ERNIE ecosystem LoRAs
  • MCP native: Integrates with Claude Desktop and other tools

For users who need rapid prompt testing, batch image generation, or AI generation integration into their own applications, the Civitai Orchestration API is currently the simplest way to call ERNIE-Image.

ERNIE-Image Team