Skip to content
Merged
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
355 changes: 355 additions & 0 deletions reference/public-http-api.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -1808,6 +1808,233 @@ paths:
'401':
description: Unauthorized
description: Marks an issue resolution as accepted by the customer.
/files:
post:
summary: Upload File
operationId: post-files
description: >-
Uploads one file as multipart/form-data. The file goes in a form part
named `file`; the stored filename comes from the part's filename and
the MIME type from the part's `Content-Type` (falling back to
`application/octet-stream`). The extension is derived from the
filename. Files are stored under the calling tenant only; uploading
the same filename twice creates two independent files with distinct
ids.
parameters:
- name: postUploadPipeline
in: query
required: false
description: >-
Optional pipeline id. Validated against the calling tenant's
pipelines before the file is stored (404, and nothing is stored,
when unknown). Once the upload is durably stored, a load of that
pipeline is started exactly as `POST /pipelines/{pipeline-id}/loads`
would — so a pipeline reading this tenant's file store picks the
new file up immediately. The upload response is unchanged: 201
with the file record, whether the load was scheduled or was a
no-op (no in-band datasets).
schema:
type: string
- name: returnRows
in: query
required: false
description: >-
Only valid together with `postUploadPipeline` (400 without it).
When true, the request waits (bounded; ~60s default) for the
triggered load to complete and the 201 body becomes
`file-upload-with-load`: the file record plus, for every dataset
in the pipeline, the rows this upload produced, read back from
the pipeline's destination. `load.status` reports what actually
happened — `written`, `declined` (no in-band datasets),
`timeout`, or `failed` — with whatever rows were readable.
Currently supported for PostgreSQL destinations only (422
otherwise).
schema:
type: boolean
default: false
requestBody:
required: true
content:
multipart/form-data:
schema:
type: object
properties:
file:
type: string
format: binary
required:
- file
responses:
'201':
description: >-
Created. Plain `file-record` without `returnRows`;
`file-upload-with-load` with it.
content:
application/json:
schema:
oneOf:
- $ref: '#/components/schemas/file-record'
- $ref: '#/components/schemas/file-upload-with-load'
'400':
description: >-
Bad Request — no form part named `file`, or `returnRows` without
`postUploadPipeline`
'401':
description: Unauthorized
'404':
description: >-
Not Found — `postUploadPipeline` names a pipeline unknown to the
calling tenant. Nothing is stored.
'413':
description: >-
Payload Too Large — the upload exceeds the deployment's byte
limit. Nothing is stored.
'422':
description: >-
Unprocessable Content — `returnRows` was requested but the named
pipeline's destination cannot serve deterministic rows. Nothing
is stored.
get:
summary: List Files
operationId: get-files
description: >-
Lists the calling tenant's files, ordered by `updatedAt` ascending
then `id` — a stable order designed for incremental consumers.
Delta-loading recipe: on the initial sync pass
`includeDeleted=true` and keep the maximum `updatedAt` seen as a
watermark; on each subsequent pass query
`updatedSince=<watermark>&includeDeleted=true` — the server returns
files uploaded or tombstoned at or after the watermark (the boundary
record repeats; dedupe by `id`), and a returned record with
`deleted: true` is a delete signal. Page forward
while `count == limit`, adding `count` to `offset`; keep
`updatedSince` fixed within one pass.
parameters:
- name: updatedSince
in: query
description: >-
Epoch milliseconds, inclusive lower bound on `updatedAt`. Pass the
last sync watermark to receive everything at or after it; the
boundary record is re-listed, and consumers dedupe by `id`
(upsert), so nothing is ever missed — including two records
sharing a millisecond.
schema:
type: integer
format: int64
- name: mimeType
in: query
description: >-
Exact MIME type match, e.g. `application/pdf`. (MIME rather than
extension — it is declared and normalized at upload; the extension
is still stored on each record.)
schema:
type: string
- name: includeDeleted
in: query
description: >-
Include tombstoned files. Delta consumers should pass `true` so
deletes propagate.
schema:
type: boolean
default: false
- name: offset
in: query
schema:
type: integer
format: int64
default: 0
- name: limit
in: query
description: Page size, clamped to 1–1000.
schema:
type: integer
default: 1000
minimum: 1
maximum: 1000
responses:
'200':
description: OK
content:
application/json:
schema:
$ref: '#/components/schemas/file-list-page'
'401':
description: Unauthorized
'/files/{fileId}':
parameters:
- name: fileId
in: path
required: true
schema:
type: string
format: uuid
get:
summary: Get File Metadata
operationId: get-files-file-id
description: >-
Returns the file record, including for tombstoned files (metadata
outlives content access). 404 when the id is unknown to the calling
tenant — including when it belongs to a different tenant.
responses:
'200':
description: OK
content:
application/json:
schema:
$ref: '#/components/schemas/file-record'
'401':
description: Unauthorized
'404':
description: Not Found
delete:
summary: Tombstone File
operationId: delete-files-file-id
description: >-
Deletes by tombstone: sets `deletedAt` and deliberately bumps
`updatedAt`, so the deletion rides the delta stream and downstream
consumers learn about it through the same `updatedSince` query as
everything else. Idempotent — deleting twice returns the record
unchanged. The stored bytes are retained, but content access answers
410 from then on.
responses:
'200':
description: OK — the updated (tombstoned) file record
content:
application/json:
schema:
$ref: '#/components/schemas/file-record'
'401':
description: Unauthorized
'404':
description: Not Found
'/files/{fileId}/content':
parameters:
- name: fileId
in: path
required: true
schema:
type: string
format: uuid
get:
summary: Get File Content
operationId: get-files-file-id-content
description: >-
Streams the file bytes with the stored MIME type as the response
`Content-Type`.
responses:
'200':
description: The bytes, exactly as uploaded
content:
'*/*':
schema:
type: string
format: binary
'401':
description: Unauthorized
'404':
description: Not Found
'410':
description: Gone — the file is tombstoned
components:
schemas:
connector:
Expand Down Expand Up @@ -3040,6 +3267,134 @@ components:
required:
- errored

file-record:
title: file-record
description: >-
A file in the tenant's file store. `updatedAt` is the delta watermark
(equal to `uploadedAt` until the file is tombstoned); `deleted` is
always present.
type: object
properties:
id:
type: string
format: uuid
readOnly: true
filename:
type: string
example: report.pdf
mimeType:
type: string
example: application/pdf
extension:
type:
- string
- 'null'
description: Derived from the filename at upload; null when the filename has none.
example: pdf
sizeBytes:
type: integer
format: int64
minimum: 0
uploadedAt:
$ref: '#/components/schemas/epoch-millis'
updatedAt:
$ref: '#/components/schemas/epoch-millis'
deletedAt:
oneOf:
- $ref: '#/components/schemas/epoch-millis'
- type: 'null'
deleted:
type: boolean
required:
- id
- filename
- mimeType
- sizeBytes
- uploadedAt
- updatedAt
- deleted
file-upload-with-load:
title: file-upload-with-load
description: >-
The 201 body of POST /files when `returnRows=true`: the stored file
plus the rows this upload produced in each of the pipeline's
datasets, read back from the destination after the triggered load.
type: object
properties:
file:
$ref: '#/components/schemas/file-record'
load:
type: object
properties:
pipelineId:
type: string
status:
type: string
enum:
- written
- declined
- timeout
- failed
description: >-
written = the load completed and rows were read; declined =
no in-band datasets, nothing to wait for; timeout/failed =
reported honestly, with whatever rows were readable.
datasets:
type: array
items:
type: object
properties:
id:
type: string
name:
type: string
rows:
type: array
description: >-
Rows keyed to this upload (destination key equals the
file id or starts with `<fileId>/`), as column-name to
value objects. Datasets fed by other sources return
empty arrays.
items:
type: object
required:
- id
- name
- rows
required:
- pipelineId
- status
- datasets
required:
- file
- load
file-list-page:
title: file-list-page
description: One page of a file listing.
type: object
properties:
files:
type: array
items:
$ref: '#/components/schemas/file-record'
count:
type: integer
description: Records in this page.
totalCount:
type: integer
format: int64
description: Records matching the query.
offset:
type: integer
format: int64
limit:
type: integer
required:
- files
- count
- totalCount
- offset
- limit
responses: {}
securitySchemes:
BearerAuth:
Expand Down
Loading