> ## Documentation Index
> Fetch the complete documentation index at: https://docs.krun.ai/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> Krun API base URL: https://api.krun.ai. Authenticate server-side with `Authorization: Bearer $KRUN_API_KEY`; never put a Krun API key in browser code.
> Prefer the Python SDK in examples: `pip install krun-ai`, `from krun import Krun`. The TypeScript SDK `@krun-ai/sdk` is not published on npm yet: do not tell users to `npm install` it.
> The only question type is `choice`. `confidence` is the top-1 minus top-2 probability margin, not the probability that the answer is correct. `choice` is null when `abstain` is true.
> Usage reports `input_tokens` only. There are no output tokens.

# Multimodal inputs

> Upcoming in Krun One V1: decide on images, documents and audio, alone or mixed with text. Not yet available on api.krun.ai.

<Warning>
  **Upcoming.** Krun One V1 multimodal inputs are not yet available on api.krun.ai. This page describes the planned
  interface so you can prepare your integration. Details can still change before release.
</Warning>

With Krun One V1, `context` can hold images, documents and audio as well as text. Every modality is **input only**:
the answer is always the same structured decision (`choice`, `noul`, `score` or [`multi`](/concepts/primitives/multi)).
Krun does not generate text, images or audio, and video is not supported.

## Text-only requests are unchanged

A string `context` keeps working exactly as today. You don't need to change existing calls:

```python theme={null}
result = client.decide(context="Customer wants to return an item.", questions={...})
```

## How it works

1. Upload each file with [`POST /v1/assets`](/api-reference/assets). You get back an `asset_id`.
2. Send `context` as a list of **content parts** that reference the assets, optionally with text parts.
3. Read the answers as usual.

Krun picks how to read each file automatically. There are no switches to set:

| Input | How Krun reads it |
| - | - |
| PDF or DOCX with a text layer | Read as text |
| Scan (scanned PDF, page image) or screenshot | OCR + vision |
| Audio | Speech transcription + acoustic features |

## Content parts

`context` is either a string (1 to 8,000 characters) or a list of 1 to 16 parts. Each part has a `type`:

| `type` | Fields | Description |
| - | - | - |
| `text` | `text` | 1 to 8,000 characters. Text parts are joined with newlines. |
| `image` | `asset_id` | An uploaded image. |
| `document` | `asset_id` | An uploaded document, or a page image such as a scan. |
| `audio` | `asset_id` | An uploaded audio clip. |

In the TypeScript SDK the field is written `assetId` (sent as `asset_id`). Every part can also carry an optional `id` (up to 100 characters), which is echoed only in errors. `asset_id` looks
like `asset_...`. No other fields are accepted: there is no `url` field and no inline file bytes.

## Image decision

<CodeGroup>
  ```bash cURL theme={null}
  ASSET_ID=$(curl -s https://api.krun.ai/v1/assets \
    -H "Authorization: Bearer $KRUN_API_KEY" \
    -H "Content-Type: image/png" \
    --data-binary @screenshot.png | jq -r .id)

  curl https://api.krun.ai/v1/decide \
    -H "Authorization: Bearer $KRUN_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "context": [
        {"type": "text", "text": "Screenshot attached by the customer."},
        {"type": "image", "asset_id": "'"$ASSET_ID"'"}
      ],
      "questions": {
        "shows_error": {"type": "noul", "instructions": "Does the screenshot show an error message?"}
      }
    }'
  ```

  ```python Python theme={null}
  from krun import Krun

  client = Krun()  # reads KRUN_API_KEY

  with open("screenshot.png", "rb") as f:
      asset = client.assets.create(f.read(), mime_type="image/png")

  result = client.decide(
      context=[
          {"type": "text", "text": "Screenshot attached by the customer."},
          {"type": "image", "asset_id": asset.id},
      ],
      questions={
          "shows_error": {"type": "noul", "instructions": "Does the screenshot show an error message?"},
      },
  )

  print(result.noul("shows_error").noul)
  ```

  ```ts TypeScript theme={null}
  import { readFile } from "node:fs/promises";
  import { Krun } from "@krun-ai/sdk";

  const krun = new Krun(); // reads KRUN_API_KEY

  const asset = await krun.assets.create(await readFile("screenshot.png"), { mimeType: "image/png" });

  const result = await krun.decide({
    context: [
      { type: "text", text: "Screenshot attached by the customer." },
      { type: "image", assetId: asset.id },
    ],
    questions: {
      shows_error: { type: "noul", instructions: "Does the screenshot show an error message?" },
    },
  });

  result.answers.shows_error.noul;
  ```
</CodeGroup>

## Document decision

Upload PDFs, DOCX, plain text, Markdown or HTML as documents. Krun reads documents with a text layer as text, and
sends scans through OCR and vision. You don't choose: the same request works for both.

<CodeGroup>
  ```bash cURL theme={null}
  ASSET_ID=$(curl -s https://api.krun.ai/v1/assets \
    -H "Authorization: Bearer $KRUN_API_KEY" \
    -H "Content-Type: application/pdf" \
    --data-binary @invoice.pdf | jq -r .id)

  curl https://api.krun.ai/v1/decide \
    -H "Authorization: Bearer $KRUN_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "context": [
        {"type": "text", "text": "Is this invoice paid?"},
        {"type": "document", "asset_id": "'"$ASSET_ID"'"}
      ],
      "questions": {
        "paid": {"type": "noul", "instructions": "Is the document marked as paid?"},
        "tags": {"type": "multi", "options": {"invoice": null, "receipt": null, "overdue": null}}
      }
    }'
  ```

  ```python Python theme={null}
  with open("invoice.pdf", "rb") as f:
      asset = client.assets.create(f.read(), mime_type="application/pdf")

  result = client.decide(
      context=[
          {"type": "text", "text": "Is this invoice paid?"},
          {"type": "document", "asset_id": asset.id},
      ],
      questions={
          "paid": {"type": "noul", "instructions": "Is the document marked as paid?"},
          "tags": {"type": "multi", "options": {"invoice": None, "receipt": None, "overdue": None}},
      },
  )
  ```

  ```ts TypeScript theme={null}
  const asset = await krun.assets.create(await readFile("invoice.pdf"), { mimeType: "application/pdf" });

  const result = await krun.decide({
    context: [
      { type: "text", text: "Is this invoice paid?" },
      { type: "document", assetId: asset.id },
    ],
    questions: {
      paid: { type: "noul", instructions: "Is the document marked as paid?" },
      tags: { type: "multi", options: { invoice: null, receipt: null, overdue: null } },
    },
  });
  ```
</CodeGroup>

A page image (PNG, JPEG or WebP), such as a scanned page, can also be sent as a `document` part.

## Audio decision

<CodeGroup>
  ```bash cURL theme={null}
  ASSET_ID=$(curl -s https://api.krun.ai/v1/assets \
    -H "Authorization: Bearer $KRUN_API_KEY" \
    -H "Content-Type: audio/wav" \
    --data-binary @voicemail.wav | jq -r .id)

  curl https://api.krun.ai/v1/decide \
    -H "Authorization: Bearer $KRUN_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "context": [{"type": "audio", "asset_id": "'"$ASSET_ID"'"}],
      "questions": {
        "department": {"type": "choice", "options": {"billing": "", "support": "", "sales": ""}},
        "urgent": {"type": "noul", "instructions": "Is the caller reporting an urgent problem?"}
      }
    }'
  ```

  ```python Python theme={null}
  with open("voicemail.wav", "rb") as f:
      asset = client.assets.create(f.read(), mime_type="audio/wav")

  result = client.decide(
      context=[{"type": "audio", "asset_id": asset.id}],
      questions={
          "department": {"type": "choice", "options": {"billing": "", "support": "", "sales": ""}},
          "urgent": {"type": "noul", "instructions": "Is the caller reporting an urgent problem?"},
      },
  )
  ```

  ```ts TypeScript theme={null}
  const asset = await krun.assets.create(await readFile("voicemail.wav"), { mimeType: "audio/wav" });

  const result = await krun.decide({
    context: [{ type: "audio", assetId: asset.id }],
    questions: {
      department: { type: "choice", options: { billing: "", support: "", sales: "" } },
      urgent: { type: "noul", instructions: "Is the caller reporting an urgent problem?" },
    },
  });
  ```
</CodeGroup>

Keep clips to 20 seconds or less: up to 30 seconds is accepted, but longer speech may be transcribed only partially.

## Mixed content

Combine text and several media parts in one `context`, in the order you want them read, up to 16 parts:

```json theme={null}
{
  "context": [
    {"type": "text", "text": "Refund request from a customer. Their message, a photo of the item and the receipt follow."},
    {"type": "text", "text": "The blender arrived with a cracked jar."},
    {"type": "image", "id": "photo", "asset_id": "asset_Hk3v9QmZ2pLwBBBBBBBBBBBB"},
    {"type": "document", "id": "receipt", "asset_id": "asset_7fQ2mZkP0aLxAAAAAAAAAAAA"}
  ],
  "questions": {
    "damaged": {"type": "noul", "instructions": "Does the photo show a damaged product?"},
    "evidence": {
      "type": "multi",
      "instructions": "Select every element present in the evidence.",
      "options": {"product_photo": null, "receipt": null, "order_number": null}
    }
  }
}
```

Several media items per request work, but their answer quality has not been validated yet. See
[Known limitations](#known-limitations).

## Limits and supported types

Per request: up to **4 images**, **2 documents** and **1 audio** part, and 1 to 16 parts in total.

| Family | MIME types | Max size |
| - | - | - |
| Image | `image/png`, `image/jpeg`, `image/webp` | 10 MB |
| Document | `application/pdf`, `application/vnd.openxmlformats-officedocument.wordprocessingml.document` (DOCX), `text/plain`, `text/markdown`, `text/html` | 25 MB, up to 20 pages |
| Audio | `audio/wav`, `audio/mpeg`, `audio/flac`, `audio/ogg` | 10 MB, up to 30 s (20 s or less recommended) |

* **Video is not supported.** A `video` part is rejected with `400 UNSUPPORTED_MODALITY`.
* Uploads are checked against the declared `Content-Type`: a PNG sent as `audio/wav` is rejected with
  `415 UNSUPPORTED_MIME_TYPE`.
* Outputs are always structured decisions. There is no text, image or audio generation.

See the [upcoming error codes](/api-reference/errors#upcoming-multimodal-error-codes).

## Privacy

* Media files, OCR text and transcripts are never logged and never returned by the API.
* Assets expire 24 hours after upload by default (`expires_at`), and you can delete them earlier with
  [`DELETE /v1/assets/{asset_id}`](/api-reference/assets#delete-an-asset).
* Assets can be used only by the project that uploaded them.

## Known limitations

* **Long audio:** clips longer than 20 seconds may be transcribed only partially.
* **Abstention on media is advisory:** treat `abstain` on image, document and audio inputs as a hint, not a validated
  guarantee. See [Abstention](/concepts/abstention).
* **Several media items per request** work, but their answer quality has not been validated.
* **Numeric comparison and hard visual negatives are weak.** Validate questions that compare numbers, or that depend on
  telling apart visually similar cases, on your own data.
