Goatlab API Docs

Working with Files

A practical guide to the two file management models in the Haufe Files API: User Files and Space Files. Understand when to use each model and how to follow the end-to-end upload and consumption flow.

Working with Files

Haufe Files supports two distinct file models depending on who owns the content and how it is shared:

User FilesSpace Files
Route prefix/v1/files/v1/spaces/{spaceId}/files
OwnershipTied to a single userTied to a tenant Space
SharingPrivate to the userPrivate (owner only) or Public (all tenant users)
Required headersAPI key + x-user-idAPI key + x-user-id + x-tenant-id
RAG scopePer-filePer-space (across all space files)

How File Processing Works

Regardless of the model, every uploaded file goes through the same processing pipeline before it can be consumed:

  1. Upload — you request a pre-signed S3 URL from the API and PUT the file directly to S3.
  2. Antivirus scan — AWS GuardDuty scans the file. Infected files are quarantined and the file status is set to FAILED.
  3. Text extraction — the appropriate processor runs based on the file type (PDF, Word, Excel, plain text / CSV / Markdown).
  4. Chunking & embedding — the extracted text is split into chunks, each one embedded via Amazon Bedrock and stored in the knowledge base.
  5. Ready — the processing status transitions to PROCESSED and the file is available for RAG or direct content retrieval.

Processing statuses

StatusMeaning
PROCESSINGPipeline is running
PROCESSEDFile is ready to use
FAILEDProcessing or antivirus check failed

Supported file types: PDF (.pdf), Word (.docx), Excel (.xlsx), PowerPoint (.pptx), plain text (.txt), CSV (.csv), Markdown (.md)

File persistence: when requesting an upload URL you pass a persist query parameter.

  • persist=true — the file is kept indefinitely.
  • persist=false — the file expires after a set period and is automatically deleted.

User Files

User files are personal — each file belongs to a single user and is only accessible by that user. Use this model when the content is user-specific and does not need to be shared.

Required headers

HeaderDescription
x-api-keyYour product API key
x-user-idThe ID of the user performing the action

Endpoints

MethodPathDescription
GET/v1/signed-urlGet a pre-signed upload URL
GET/v1/filesList all files for the user (paginated)
GET/v1/files/{fileId}Retrieve the processed text content of a file
GET/v1/files/{fileId}/statusProcessing status of a file
GET/v1/files/{fileId}/ragSemantic search within a single file
DELETE/v1/files/{fileId}Delete a file

Basic use-case flow

Below is the minimal sequence of API calls to upload a file and use it for RAG.

Step 1 — Request an upload URL
────────────────────────────────
GET /v1/signed-url?filename=report.pdf&persist=true
Headers:
  x-api-key: <your-api-key>
  x-user-id: <user-id>

Response:
{
  "url": "https://s3.amazonaws.com/...",   ← pre-signed PUT URL
  "file_id": "3f2e1d...",                  ← keep this ID
  "expires_in": 600,
  "file_limits": {
    "max_size": 10485760,
    "allowed_types": ["application/pdf", "text/plain", ...]
  }
}
Step 2 — Upload the file directly to S3
────────────────────────────────────────
PUT <url from step 1>
Content-Type: application/pdf   ← must match the file type
Body: <binary file content>

Response: 200 OK (from S3, no body)
Step 3 — Check processing status
───────────────────────────────────────────
GET /v1/files/{fileId}/status
Headers:
  x-api-key: <your-api-key>
  x-user-id: <user-id>

Response:
{
  "processingStatus": "PROCESSED",   ← poll until this is "PROCESSED" or "FAILED"
  "steps": [
    { "stepName": "FirstStep",         "status": "SUCCEEDED" },
    { "stepName": "ProcessorSelector", "status": "SUCCEEDED" },
    { "stepName": "PdfProcessor",      "status": "SUCCEEDED" },
    { "stepName": "ChunkFile",         "status": "SUCCEEDED" }
  ],
  "warnings": { "length": false, "empty": false }
}
Step 4 — Query the file with RAG
──────────────────────────────────
GET /v1/files/{fileId}/rag?query=What+are+the+main+findings&limit=5
Headers:
  x-api-key: <your-api-key>
  x-user-id: <user-id>

Response:
[
  {
    "content": "The analysis shows ...",
    "metadata": { "chunk_id": "...", "file_id": "..." },
    "score": 0.94
  },
  ...
]
Step 5 (optional) — Retrieve the full processed content
────────────────────────────────────────────────────────
GET /v1/files/{fileId}
Headers:
  x-api-key: <your-api-key>
  x-user-id: <user-id>

Response: text/plain — the extracted and processed text of the file
Step 6 (optional) — Delete the file
─────────────────────────────────────
DELETE /v1/files/{fileId}
Headers:
  x-api-key: <your-api-key>
  x-user-id: <user-id>

Response: { "message": "File <fileId> deleted successfully." }

Tip: if you uploaded with persist=false, the file will be automatically cleaned up after it expires — you do not need to call DELETE explicitly.


Space Files

Space files belong to a Space — a tenant-scoped container that groups files under a shared context. Spaces have a visibility setting that controls who can access them:

  • PRIVATE (default) — only the space owner can read, upload, or delete files in the space.
  • PUBLIC — any user within the same tenant can access the space and its files.

Use this model when you want to perform RAG queries over a collection of documents, or when content needs to be shared across users within a tenant.

A Space must exist before files can be uploaded to it. Refer to the Spaces endpoints (POST /v1/spaces) to create one.

Required headers

HeaderDescription
x-api-keyYour product API key
x-user-idThe ID of the user performing the action
x-tenant-idThe ID of the tenant the space belongs to

Endpoints

MethodPathDescription
GET/v1/spaces/{spaceId}/files/upload-urlGet a pre-signed upload URL for the space
GET/v1/spaces/{spaceId}/filesList all files in the space (paginated)
GET/v1/spaces/{spaceId}/files/{fileId}Retrieve the processed text content of a file
GET/v1/spaces/{spaceId}/files/{fileId}/statusPoll the processing status of a file
GET/v1/spaces/{spaceId}/ragSemantic search across all files in the space
DELETE/v1/spaces/{spaceId}/files/{fileId}Delete a file from the space

Basic use-case flow

Step 0 — Make sure you have a Space
──────────────────────────────────────
POST /v1/spaces
Headers:
  x-api-key:   <your-api-key>
  x-user-id:   <user-id>
  x-tenant-id: <tenant-id>
Body:
{
  "name": "My Knowledge Base",
  "visibility": "PRIVATE"   // default; use "PUBLIC" to share with all users in the tenant
}

Response:
{
  "id": "space-uuid",
  "name": "My Knowledge Base",
  ...
}
Step 1 — Request an upload URL for the space
─────────────────────────────────────────────
GET /v1/spaces/{spaceId}/files/upload-url?filename=guidelines.pdf&persist=true
Headers:
  x-api-key:   <your-api-key>
  x-user-id:   <user-id>
  x-tenant-id: <tenant-id>

Response:
{
  "url":      "https://s3.amazonaws.com/...",
  "file_id":  "8a1c3f...",
  "expires_in": 600,
  "file_limits": {
    "max_size": 10485760,
    "allowed_types": ["application/pdf", "text/plain", ...]
  }
}
Step 2 — Upload the file directly to S3
────────────────────────────────────────
PUT <url from step 1>
Content-Type: application/pdf
Body: <binary file content>

Response: 200 OK (from S3, no body)
Step 3 — Check processing status
───────────────────────────────────────────
GET /v1/spaces/{spaceId}/files/{fileId}/status
Headers:
  x-api-key:   <your-api-key>
  x-user-id:   <user-id>
  x-tenant-id: <tenant-id>

Response:
{
  "processingStatus": "PROCESSED",
  "steps": [
    { "stepName": "FirstStep",         "status": "SUCCEEDED" },
    { "stepName": "ProcessorSelector", "status": "SUCCEEDED" },
    { "stepName": "PdfProcessor",      "status": "SUCCEEDED" },
    { "stepName": "ChunkFile",         "status": "SUCCEEDED" }
  ],
  "warnings": { "length": false, "empty": false }
}
Step 4 — Query the entire space with RAG
──────────────────────────────────────────
GET /v1/spaces/{spaceId}/rag?query=What+are+the+holiday+policies&limit=5
Headers:
  x-api-key:   <your-api-key>
  x-user-id:   <user-id>
  x-tenant-id: <tenant-id>

Response:
[
  {
    "content": "Employees are entitled to ...",
    "metadata": {
      "chunk_id": "...",
      "file_id":  "...",
      "space_id": "..."
    },
    "score": 0.97
  },
  ...
]

The Space RAG endpoint searches across all processed files in the space in a single call. You do not need to query each file individually.

Step 5 (optional) — List files in the space
─────────────────────────────────────────────
GET /v1/spaces/{spaceId}/files?page=1&limit=20&sortBy=createdAt&sortOrder=desc
Headers:
  x-api-key:   <your-api-key>
  x-user-id:   <user-id>
  x-tenant-id: <tenant-id>

Response:
{
  "data": [
    {
      "id": "8a1c3f...",
      "originalName": "guidelines.pdf",
      "processingStatus": "PROCESSED",
      "size": 204800,
      "createdAt": "2026-03-19T10:00:00Z",
      ...
    }
  ],
  "total": 1,
  "page": 1,
  "limit": 20
}
Step 6 (optional) — Delete a file from the space
──────────────────────────────────────────────────
DELETE /v1/spaces/{spaceId}/files/{fileId}
Headers:
  x-api-key:   <your-api-key>
  x-user-id:   <user-id>
  x-tenant-id: <tenant-id>

Response: { "message": "File deleted successfully." }

Choosing the Right Model

ScenarioRecommended model
A user uploads their own private documentsUser Files
Multiple documents are shared across a team / tenantSpace Files
You need RAG over a single documentUser Files — use /v1/files/{fileId}/rag
You need RAG over a collection of documentsSpace Files — use /v1/spaces/{spaceId}/rag
Content is temporary (used once, then discarded)Either model with persist=false
Content is permanent (knowledge base, product data)Either model with persist=true

Haufe Files API Docs

Haufe Files API documentation provides detailed information about the API endpoints, request and response formats, and authentication methods. It is designed to help developers integrate with the Haufe Files platform effectively. The whole OpenApi specification can be found [here](/api/openapi). ## API Key Modes This API operates in two distinct modes based on the type of API key used: --- # **Product Mode (General Usage)** Designed for Haufe Products leveraging BYOC (Bring Your Own Content) capabilities to manage customer content through Haufe Files. - Non-Scoped API Keys: - Used by HaufeShop products (e.g., Copilots) - User and tenant management provided by Control Plane - Can access our UI if the IDs match those in Control Plane - Scoped API Keys: - Used by products like HR Assistant or KaaS - Can only access data within their own scope - Automatically provisions tenants and users when needed - User and tenant IDs are prefixed with the scope (e.g., `{scope}:{externalId}`) --- # **Gen AI Mode** A specialized API key type for generative AI tools requiring access to uploaded data for LLM integration. - **Exclusive to HAI (Haufe AI)** - Can read data from all scopes with the appropriate IDs - Limited action set (read-focused capabilities) - Permissions and features will evolve as the product develops --- Feel free to contact the following people in case you have any questions: - Ioana Lefter (Product Owner) - Pavel Khralovich (Tech Lead) - Raul de San Clemente (Software Engineer) - Alejandro Cano (Software Engineer) - Giulia Mainiero (Software Engineer)

Check API health GET

Next Page

On this page