Skip to main content

Documents Automation

Availability: folder discovery, original previews, typed template authoring and scan monitoring are implemented in the development build. Your deployment needs the matching backend and UI. OCR, extraction execution, template activation, result review and publication are still pending; saving a template does not run extraction.

The sidebar entry is Documents Automation, immediately after Connectors and before Settings & Monitoring. The workspace has Documents, Templates and Monitoring tabs and follows the app's light/dark theme.

Prepare and enable the workspace​

  1. Ask your deployment administrator to deploy the matching backend and UI and confirm registry startup completed. The backend adds the document tables through its registry schema initialization. This is not a tenant extraction toggle.
  2. Sign in and select the correct tenant and project. An editor, admin or super-admin can attach folders, scan and save templates; viewers can browse authorized inventories, originals and saved versions.
  3. In Connectors, configure an existing file-store adapter for a local POSIX folder, S3 or an on-prem folder. Grant that store to the selected project and its intended users. A database connector alone is not a folder adapter. Use least-privilege folder/bucket permissions.
  4. For on-prem access, upgrade to connector 1.2.0 or later using the bundle downloaded from the updated backend. Keep all bundle files, including document_files.py, together. Configure FILES_DIR for the approved folder and keep the connector online. Safe local document reads currently require POSIX; Windows document access is not certified by this release.
  5. Open Documents Automation → Documents, choose Connect folder, select the existing adapter, enter a display name and a relative subfolder, then save. Leave the subfolder empty to use the configured root. Absolute paths and parent traversal are rejected.
  6. Select the folder and choose Scan. Review the inventory state before opening documents. A partial scan is not a complete inventory.

If the entry is missing, first check the deployed UI version. If no adapter is listed, check tenant ownership, project assignment, permissions and adapter type. Shared stores from another tenant are not implicitly granted to this workspace.

Browse documents and create a template​

  1. Select a scanned folder. The document table shows name/path, type, size and discovery state. Pagination reads a pinned scan snapshot.
  2. Choose View on a discovered document. PDF, PNG and JPEG originals are previewed after bounded reads and file-signature checks. This preview does not invoke an extraction model.
  3. Choose Create template, or open Templates. Enter a template name and an actual applicability/classification rule describing the document family.
  4. Define output fields with their names, types and required flags. Review the output-column preview. Use the advanced JSON editor for runtime inputs, repeated tables and richer supported contracts.
  5. Validate and save. Each save creates an immutable revision; an existing revision is not overwritten. Open version history to inspect or export the saved contract.

Typed contracts support identifiers/text, dates, decimals, booleans and configured enumerations. Validation checks references, types and supported rules. The current editor defines what should be extracted; it does not yet populate an output table. Automatic acceptance and template activation are disabled.

Monitor folder scans​

Open Monitoring to review folder, adapter, scan state, observed count, last scan time and reported issues. Editors can rescan; viewers can inspect status. This tab currently monitors discovery, not extraction jobs.

State or issueWhat to do
Not scanned / empty inventoryRun a scan and confirm the selected subfolder contains supported files.
Partial scanResolve reported access, depth or inventory-limit issues; do not treat the count as complete.
Adapter unavailable / offlineCheck the adapter configuration and connector connectivity, then rescan. A previously usable snapshot may remain visible.
Unsupported / blocked fileCheck supported type and path policy. Symlinks and special files are blocked for safe local reads.
Too largeReview the configured transfer limit with the deployment administrator.
Original changed since scanRescan before viewing the changed original.

Scans are synchronous and capped. Complete large-folder paging, deletion lifecycle tracking and retained immutable originals are not implemented yet.

Deployment limits and security​

Deployment administrators set max_file_bytes, inventory_limit, inventory_depth, page_size, max_template_bytes and result_retention_seconds in the document_automation section of config.yaml. DOCUMENT_MAX_FILE_BYTES overrides the binary limit through the environment. The default binary limit is 20 MB; confirm the active backend capability response rather than assuming a limit.

GET /api/document-automation/capabilities reports actual transfer and execution capabilities for an authenticated user. It currently reports extraction, automatic acceptance and template activation as disabled. There is no supported switch to enable those unfinished features.

Folder bindings, snapshots, previews and template versions are scoped to tenant/project. Original access is rechecked after the adapter read. On-prem binary job results are encrypted transiently and cleared after consumption; unconsumed stale results are swept on subsequent connector polls. Document bytes can pass through the backend for preview, so deployment and residency policy still matter.

Documents Automation does not yet provide Windows user impersonation or downstream per-user delegation. Review Enterprise identity & Active Directory before promising AD permission inheritance.