Skip to main content
Upcoming. Krun One V1 multimodal inputs are not yet available on api.krun.ai. This page describes the planned interface so you can prepare your integration. Details can still change before release.
With Krun One V1, context can hold images, documents and audio as well as text. Every modality is input only: the answer is always the same structured decision (choice, noul, score or multi). Krun does not generate text, images or audio, and video is not supported.

Text-only requests are unchanged

A string context keeps working exactly as today. You don’t need to change existing calls:

How it works

  1. Upload each file with POST /v1/assets. You get back an asset_id.
  2. Send context as a list of content parts that reference the assets, optionally with text parts.
  3. Read the answers as usual.
Krun picks how to read each file automatically. There are no switches to set:

Content parts

context is either a string (1 to 8,000 characters) or a list of 1 to 16 parts. Each part has a type: In the TypeScript SDK the field is written assetId (sent as asset_id). Every part can also carry an optional id (up to 100 characters), which is echoed only in errors. asset_id looks like asset_.... No other fields are accepted: there is no url field and no inline file bytes.

Image decision

Document decision

Upload PDFs, DOCX, plain text, Markdown or HTML as documents. Krun reads documents with a text layer as text, and sends scans through OCR and vision. You don’t choose: the same request works for both.
A page image (PNG, JPEG or WebP), such as a scanned page, can also be sent as a document part.

Audio decision

Keep clips to 20 seconds or less: up to 30 seconds is accepted, but longer speech may be transcribed only partially.

Mixed content

Combine text and several media parts in one context, in the order you want them read, up to 16 parts:
Several media items per request work, but their answer quality has not been validated yet. See Known limitations.

Limits and supported types

Per request: up to 4 images, 2 documents and 1 audio part, and 1 to 16 parts in total.
  • Video is not supported. A video part is rejected with 400 UNSUPPORTED_MODALITY.
  • Uploads are checked against the declared Content-Type: a PNG sent as audio/wav is rejected with 415 UNSUPPORTED_MIME_TYPE.
  • Outputs are always structured decisions. There is no text, image or audio generation.
See the upcoming error codes.

Privacy

  • Media files, OCR text and transcripts are never logged and never returned by the API.
  • Assets expire 24 hours after upload by default (expires_at), and you can delete them earlier with DELETE /v1/assets/{asset_id}.
  • Assets can be used only by the project that uploaded them.

Known limitations

  • Long audio: clips longer than 20 seconds may be transcribed only partially.
  • Abstention on media is advisory: treat abstain on image, document and audio inputs as a hint, not a validated guarantee. See Abstention.
  • Several media items per request work, but their answer quality has not been validated.
  • Numeric comparison and hard visual negatives are weak. Validate questions that compare numbers, or that depend on telling apart visually similar cases, on your own data.