> ## Documentation Index
> Fetch the complete documentation index at: https://docs.lendflow.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Document Data Extraction

## Document Data Extraction

The **Document Data Extraction** Workflow Builder block sends one business document to Lendflow's extraction service and stores the categorized, extracted result. The current block represents `arsen_ai_document_data_extraction`.

## Requirements

| Application field | Requirement | Notes |
| - | - | - |
| Newly uploaded document | Conditional | Select exactly one source; PDF only and subject to Lendflow's configured file-size limit. |
| Existing application document | Conditional | Select exactly one source; the stored PDF must be resolvable. |

## API flow

Use the dedicated document-extraction endpoint, not the general enrichment endpoint:

1. Send `POST /api/applications/{application_id}/document/extract_data`.
2. Submit either multipart `file` or `selected_file_id`.
3. A successful request returns HTTP `201` with an empty body and queues extraction.
4. Poll [Get Commercial Data](/api-reference/workflow-management/get-commercial-data) with `services[]=arsen_ai_document_data_extraction`.

```text theme={"system"}
POST /api/applications/{application_id}/document/extract_data
Authorization: Bearer {token}
Content-Type: multipart/form-data

file=@sanitized-document.pdf
```

## Data Orchestration availability and flow

The block is available in the **Data Extraction** group for business entities. Configure it with the path of an attached or uploaded PDF. Each execution processes one file and appends a stored extraction record; it does not merge multiple documents into one response.

## What the service returns

| Result area | Response path | Meaning |
| - | - | - |
| Extraction records | `data.commercial_data.document_data_extraction` | A collection of stored results for the business. |
| Document category | Each record's `document_type` | `categorization_result.category_group`, or `UNKNOWN` when that value is absent. |
| Provider payload | Each record's `response` | The extraction service response, stored without a fixed public field-level schema. |
| Provider status | Each record's `status_code` | HTTP status returned by the extraction service. |

## Representative response

The extraction schema depends on the uploaded document. This example shows only fields confirmed by the storage flow.

```json theme={"system"}
{
  "data": {
    "commercial_data": {
      "document_data_extraction": [{
        "document_type": "TAX_RETURN",
        "status_code": 200,
        "response": {
          "categorization_result": {
            "category_group": "TAX_RETURN"
          }
        }
      }]
    }
  }
}
```

## Field meanings

| Field | Type | Meaning |
| - | - | - |
| `document_type` | String | Stored category group. Falls back to `UNKNOWN` if the provider omits it. |
| `status_code` | Integer | Provider HTTP response code. |
| `response` | Object | Provider-defined categorization and extraction payload. Fields vary by recognized document type. |
| `categorization_result.category_group` | String or null | Provider category used to populate `document_type`. |

Do not assume a tax-return, bank-statement, or other document schema unless those fields are present in the returned `response`.

## Errors and statuses

| Signal | Meaning |
| - | - |
| HTTP `201` | The extraction was queued; the response body does not contain the extracted data. |
| HTTP `422` with `file` | The upload is missing, too large, or not a PDF. |
| HTTP `422` with `selected_file_id` | The ID is missing, does not exist, or does not resolve to a stored file. |
| HTTP `403` | The caller cannot create a file for the application. |
| Stored `status_code` of `400` or greater | The extraction provider rejected or failed the document. |
| `document_type: "UNKNOWN"` | The response did not contain a category group; it does not by itself identify the cause. |

## FAQ

<AccordionGroup>
  <Accordion title="Can I submit an image directly?">
    No. The dedicated extraction request currently validates `mimes:pdf`.
  </Accordion>

  <Accordion title="Can I extract an existing application file?">
    Yes. Send its valid `selected_file_id` instead of uploading `file`.
  </Accordion>

  <Accordion title="Why does this guide not list every extracted field?">
    The provider response varies by document type, and the current Lendflow code does not define one stable public extraction schema. Inspect the returned `response` instead of assuming fields.
  </Accordion>
</AccordionGroup>
