## Add a training file

**post** `/v2/training_data/files`

Adds a document the AI agent can learn from. Provide exactly one source:

- **Inline text** — `{ "name": "refund-policy.md", "content": "..." }`. Markdown or plain text is stored as-is and indexing is attempted before the response returns; check `indexStatus` (`status` is `completed`). Best for content you generate yourself, such as an SOP compiled from resolved conversations.
- **Remote file** — `{ "url": "https://..." }` (optional `name`). The file (PDF, DOCX, PPTX, XLSX, TXT, HTML, …; max 50 MB) is downloaded and converted to text in the background. The response returns immediately with `status: "processing"`; poll the file until it is `completed` or `failed`.
- **Multipart upload** — send `multipart/form-data` with the document in a `file` part and, optionally, a JSON `data` field such as `{"name": "..."}`. Converted in the background like a remote file.

**Keeping files in sync (create-or-update).** Pass your own `externalId` (unique per workspace) and the same request becomes idempotent: a new ID creates the file (`201`), a known ID with different content updates that file in place and re-indexes it (`200`), and a known ID with identical content is a no-op (`200`, no embedding work). Add a `source` label to group everything one pipeline owns, then list by `source` to find entries to remove. Every response carries `contentHash` so you can diff a whole library from the list endpoint.

Without `externalId`, uploading a file whose name and bytes match an existing file returns the existing file with status `200` instead of creating a duplicate. Multipart uploads over 50 MB are rejected with `413`.

### Header Parameters

- `"Featurebase-Version": optional "2026-08-19.orbit" or "2026-01-01.nova" or "2025-12-12.clover"`

  - `"2026-08-19.orbit"`

  - `"2026-01-01.nova"`

  - `"2025-12-12.clover"`

### Body Parameters

- `content: optional string`

  Text or markdown to train on. Stored as-is; indexing is attempted before returning. Check indexStatus. Cannot be combined with `url`.

- `externalId: optional string`

  Your own identifier for this entry (unique per workspace). When provided, the request creates the entry if the ID is new and otherwise updates the existing entry in place — an unchanged payload is a no-op. Use it to keep the AI agent in sync with a system of record without tracking Featurebase IDs.

- `name: optional string`

  Display name of the file. Required with `content`; optional with `url` (defaults to the downloaded file name) or a multipart upload (defaults to the uploaded file name).

- `source: optional string`

  Free-form label for the pipeline or system this entry came from. Filter lists by it to review or sweep everything from one source.

- `url: optional string`

  Public http(s) URL of a document to download and convert (PDF, DOCX, TXT, HTML, …; max 50 MB). Conversion runs in the background — poll the file until `status` is `completed`. Cannot be combined with `content`.

### Returns

- `TrainingFile object { id, characterCount, contentHash, 15 more }`

  - `id: string`

    Training file ID

  - `characterCount: number`

    Characters of extracted text

  - `contentHash: string`

    SHA-256 (hex) of the stored source: for inline text and content edits the trimmed text, for uploads and URLs the original bytes. Compare it with your own hash to detect changes without downloading `content`.

  - `contentPreview: string`

    First 200 characters of the extracted text

  - `createdAt: string`

    ISO timestamp of the upload

  - `externalId: string`

    Your own identifier for this file, if one was set

  - `fileSize: number`

    Size of the originally uploaded source in bytes (unchanged by content edits)

  - `indexStatus: "pending" or "indexed" or "failed"`

    Whether the content is searchable by the AI agent. `pending` while indexing, `indexed` when the AI agent can use it, `failed` if indexing failed (the content is stored and indexing is retried hourly).

    - `"pending"`

    - `"indexed"`

    - `"failed"`

  - `mimeType: string`

    MIME type of the originally uploaded source

  - `name: string`

    File name shown in the dashboard

  - `object: "training_file"`

    Object type identifier

    - `"training_file"`

  - `processingError: string`

    Conversion error message when `status` is `failed`

  - `source: string`

    Pipeline label, if one was set

  - `status: "processing" or "completed" or "failed"`

    `processing` while the file is being converted to text, `completed` once the content is stored, `failed` if conversion failed (see `processingError`).

    - `"processing"`

    - `"completed"`

    - `"failed"`

  - `updatedAt: string`

    ISO timestamp of the last change

  - `wordCount: number`

    Words of extracted text

  - `content: optional string`

    Full extracted text (markdown). Returned when retrieving or updating a single file; omitted from list and create responses.

  - `outcome: optional "created" or "updated" or "unchanged"`

    On create and update responses: `created`, `updated` (same `externalId` or same name and bytes, content replaced and re-indexed), or `unchanged` (identical content, nothing done). Absent on reads.

    - `"created"`

    - `"updated"`

    - `"unchanged"`

### Example

```http
curl https://do.featurebase.app/v2/training_data/files \
    -X POST \
    -H "Authorization: Bearer $FEATUREBASE_API_KEY"
```

#### Response

```json
{
  "id": "67ec1234abcd5678ef901234",
  "characterCount": 12840,
  "contentHash": "9f86d081884c7d659a2feaa0c55ad015a3bf4f1b2b0b822cd15d6c15b0f00a08",
  "contentPreview": "# Refund policy\n\nCustomers can request a refund within 30 days…",
  "createdAt": "2026-09-03T10:15:00.000Z",
  "externalId": "kb-article-1842",
  "fileSize": 48213,
  "indexStatus": "indexed",
  "mimeType": "application/pdf",
  "name": "refund-policy.pdf",
  "object": "training_file",
  "processingError": null,
  "source": "resolved-conversations",
  "status": "completed",
  "updatedAt": "2026-09-03T10:15:30.000Z",
  "wordCount": 2130,
  "content": "# Refund policy\n\nCustomers can request a refund within 30 days of purchase…",
  "outcome": "created"
}
```
