PDF (Portable Document Format)
PDF is a fixed-layout document format designed to preserve the appearance of text, tables, images, and pages across systems.
Financial PDFs can carry statements, invoices, reports, contracts, tax documents, and exports. The file is a document container, not automatically structured financial data.
A PDF may contain readable text, embedded tables, or only scanned images. Extracted values need review because layout, OCR, table structure, and source quality can change what a system can interpret reliably.
Move from source document to reviewed input
Extraction can propose structure, but review remains the boundary before financial records are written.
- 01Source PDF
A fixed-layout financial document with its own source context
- 02Extraction
Readable text and fields become suggestions
- 03Review
A person checks the source, values, category, and destination
- 04Structured input
Only approved supported items become records
What is a PDF?
Portable Document Format, or PDF, preserves the visual layout of a document so it can be opened and shared consistently across devices. A PDF can contain selectable text, tables, vector elements, embedded images, or pages that are only scans.
In finance, PDFs commonly carry bank statements, invoices, receipts, payroll reports, contracts, tax documents, management reports, and exported financial statements. The format says how the document is packaged and displayed. It does not establish that every value is correct, current, complete, or ready for calculation.
Structured data has explicit fields and relationships that software can process consistently. A PDF usually preserves presentation instead. Turning it into structured input may require text parsing, table interpretation, OCR for image-only pages, and human validation against the source.
Why is a PDF not structured financial data?
A person can often understand a statement because headings, columns, spacing, and page order provide context. Software must infer that structure unless the file contains reliable tagged data. Repeated headers, wrapped descriptions, multi-page tables, negative-number formats, and merged cells can all affect extraction.
Text-based PDF versus scanned PDF
A text-based PDF contains readable text that can be parsed directly. A scanned PDF is primarily an image and generally requires optical character recognition before its contents can be interpreted. OCR can introduce character, decimal, date, or column errors, so the extraction method and its limitations should remain visible.
What should be reviewed after extraction?
Source context
Confirm the document type, period, entity, account, currency, and whether every relevant page is present.
Values and signs
Check amounts, dates, debit or credit direction, totals, and any opening or closing balance.
Classification and destination
An extracted label is a suggestion until the category and destination record are approved.
Why it matters
Financial documents often arrive in PDF because the format is easy to distribute and preserves an official-looking layout. That convenience can hide a data-quality boundary. A visually polished statement can still be incomplete, scanned poorly, or interpreted incorrectly.
Keeping the source document, extracted suggestion, review decision, and approved record distinct makes errors easier to catch and creates a clearer audit trail for later planning.
What goes into it
- A supported text-based PDF with readable source text
- Document type, reporting period, entity, account, and currency context
- Extracted fields linked back to a source excerpt
- Human review before supported records are written
Illustrative document review
A company uploads a text-based bank-statement PDF. The extractor proposes transaction dates, descriptions, amounts, and a closing balance. One table row wrapped across two lines, so the description needs correction. The user checks the source, corrects the suggestion, routes supported items, and approves only the records that match the document.
How RunwayCal helps
RunwayCal Upload accepts one supported text-based PDF or CSV at a time. It reads available document text, uses AI to suggest supported financial items, and keeps confidence, source excerpts, review flags, and routing choices visible before anything is saved.
Image-only PDFs are not processed by the current workflow because it does not perform image OCR. Extracted items remain suggestions until a user edits, routes, approves, or skips them. The source file itself is not retained after extraction.
Common mistakes
- 1Treating a PDF as structured data merely because the layout looks tabular.
- 2Assuming every PDF contains readable text or that OCR is always available and accurate.
- 3Writing extracted suggestions without checking the source, period, currency, signs, and destination.
- 4Describing a document upload as a live bank connection or autonomous bookkeeping workflow.
Get the Financial Clarity Newsletter
Practical tips on cash flow, runway, and financial decisions for founders, business owners, CFOs, investors, and board members. Free, weekly, no spam.
Extract the details. Keep the write under your control.
Review every supported PDF suggestion against its source before it becomes financial data.
Explore Upload