Documents Automation
Availability: folder discovery, original previews, typed template authoring and scan monitoring are implemented in the development build. Your deployment needs the matching backend and UI. OCR, extraction execution, template activation, result review and publication are still pending; saving a template does not run extraction.
The sidebar entry is Documents Automation, immediately after Connectors and before Settings & Monitoring. The workspace has Documents, Templates and Monitoring tabs and follows the app's light/dark theme.
Prepare and enable the workspace
- Ask your deployment administrator to deploy the matching backend and UI and confirm registry startup completed. The backend adds the document tables through its registry schema initialization. This is not a tenant extraction toggle.
- Sign in and select the correct tenant and project. An editor, admin or super-admin can attach folders, scan and save templates; viewers can browse authorized inventories, originals and saved versions.
- In Connectors, configure an existing file-store adapter for a local POSIX folder, S3 or an on-prem folder. Grant that store to the selected project and its intended users. A database connector alone is not a folder adapter. Use least-privilege folder/bucket permissions.
- For on-prem access, upgrade to connector 1.2.0 or later using the bundle downloaded from the updated backend. Keep all bundle files, including
document_files.py, together. ConfigureFILES_DIRfor the approved folder and keep the connector online. Safe local document reads currently require POSIX; Windows document access is not certified by this release. - Open Documents Automation → Documents, choose Connect folder, select the existing adapter, enter a display name and a relative subfolder, then save. Leave the subfolder empty to use the configured root. Absolute paths and parent traversal are rejected.
- Select the folder and choose Scan. Review the inventory state before opening documents. A partial scan is not a complete inventory.
If the entry is missing, first check the deployed UI version. If no adapter is listed, check tenant ownership, project assignment, permissions and adapter type. Shared stores from another tenant are not implicitly granted to this workspace.
Browse documents and create a template
- Select a scanned folder. The document table shows name/path, type, size and discovery state. Pagination reads a pinned scan snapshot.
- Choose View on a discovered document. PDF, PNG and JPEG originals are previewed after bounded reads and file-signature checks. This preview does not invoke an extraction model.
- Choose Create template, or open Templates. Enter a template name and an actual applicability/classification rule describing the document family.
- Define output fields with their names, types and required flags. Review the output-column preview. Use the advanced JSON editor for runtime inputs, repeated tables and richer supported contracts.
- Validate and save. Each save creates an immutable revision; an existing revision is not overwritten. Open version history to inspect or export the saved contract.
Typed contracts support identifiers/text, dates, decimals, booleans and configured enumerations. Validation checks references, types and supported rules. The current editor defines what should be extracted; it does not yet populate an output table. Automatic acceptance and template activation are disabled.
Monitor folder scans
Open Monitoring to review folder, adapter, scan state, observed count, last scan time and reported issues. Editors can rescan; viewers can inspect status. This tab currently monitors discovery, not extraction jobs.
| State or issue | What to do |
|---|---|
| Not scanned / empty inventory | Run a scan and confirm the selected subfolder contains supported files. |
| Partial scan | Resolve reported access, depth or inventory-limit issues; do not treat the count as complete. |
| Adapter unavailable / offline | Check the adapter configuration and connector connectivity, then rescan. A previously usable snapshot may remain visible. |
| Unsupported / blocked file | Check supported type and path policy. Symlinks and special files are blocked for safe local reads. |
| Too large | Review the configured transfer limit with the deployment administrator. |
| Original changed since scan | Rescan before viewing the changed original. |
Scans are synchronous and capped. Complete large-folder paging, deletion lifecycle tracking and retained immutable originals are not implemented yet.
Deployment limits and security
Deployment administrators set max_file_bytes, inventory_limit, inventory_depth, page_size, max_template_bytes and result_retention_seconds in the document_automation section of config.yaml. DOCUMENT_MAX_FILE_BYTES overrides the binary limit through the environment. The default binary limit is 20 MB; confirm the active backend capability response rather than assuming a limit.
GET /api/document-automation/capabilities reports actual transfer and execution capabilities for an authenticated user. It currently reports extraction, automatic acceptance and template activation as disabled. There is no supported switch to enable those unfinished features.
Folder bindings, snapshots, previews and template versions are scoped to tenant/project. Original access is rechecked after the adapter read. On-prem binary job results are encrypted transiently and cleared after consumption; unconsumed stale results are swept on subsequent connector polls. Document bytes can pass through the backend for preview, so deployment and residency policy still matter.
Documents Automation does not yet provide Windows user impersonation or downstream per-user delegation. Review Enterprise identity & Active Directory before promising AD permission inheritance.