1. Upload JSONL
Each line is an OpenAI batch request: custom_id, method, url, and a chat body.
<batchrate>>
Upload a JSONL file of chat requests. Results come back within a 24 hour window, usually within hours. The API matches the OpenAI Batch API, so the official SDK works after you change base_url. Rates are loaded from config and listed next to reference prices on the pricing page.
Each line is an OpenAI batch request: custom_id, method, url, and a chat body.
Requests run on interruptible GPUs. If a machine disappears, its lease expires and another worker picks up the same chunk.
Completed lines come back as an output JSONL file. Failures land in an error file. You pay for tokens actually generated.
from openai import OpenAI
client = OpenAI(
api_key="br_live_...",
base_url="https://batchrate.ai/v1",
)
batch_file = client.files.create(
file=open("batch.jsonl", "rb"),
purpose="batch",
)
batch = client.batches.create(
input_file_id=batch_file.id,
endpoint="/v1/chat/completions",
completion_window="24h",
)
print(batch.id, batch.status)
A sample file is at /sample-batch.jsonl. The first model is Qwen3.6-35B-A3B FP8, served as qwen3.6-35b-a3b.
A verified account receives a one-time credit of 10,000,000 tokens, shared by input and output, and used before paid credits. After that, add credits with Stripe Checkout. New batches are rejected when free tokens plus the available balance cannot cover the estimate.
Estimate a batchStage 1 accepts and queues work. A separate worker, running later on interruptible Vast.ai capacity, claims chunks through a private ops API. Lost workers do not lose the batch.
API reference