Uploaded PDFs, Word files, spreadsheets and scanned images are parsed into searchable, project-isolated chunks and answered over with retrieval-grounded chat.

Where it started
Due-diligence material arrives as whatever the counterparty happens to send: born-digital PDFs, Word documents, spreadsheets, and scans that are images of text. Reviewers were reading everything manually. Any assistant had to keep each project's documents strictly separate — a question asked inside one engagement must never retrieve from another.
Mixed-format ingestion
PDF, Word, Excel and scanned images all enter the same pipeline, with OCR for pages that are images of text.
Project isolation
Each engagement's documents are sealed from every other, enforced in retrieval rather than filtered afterwards.
Grounded chat
Questions are answered over the indexed corpus, with answers traceable back to the source material.
Background processing
Parsing, embedding and indexing run as queued jobs, so a large upload never blocks the interface.
Role-based access
Guard-based permissions on every endpoint, applied server-side.
Search across a corpus
Reviewers query a body of documents instead of reading each one end to end.
- FastAPI service with PostgreSQL, using pgvector alongside the relational data rather than a separate vector store to maintain.
- Retrieval pipeline built on LangChain, with project scoping applied at the retrieval layer.
- Celery workers behind Redis for parsing, embedding and indexing, keeping heavy work off the request path.
- JWT authentication with role and guard checks enforced in the API, not the client.
- React 19 with TypeScript, Vite and MUI in an Nx monorepo shared with the other front ends.
- Containerised and deployed through an automated pipeline.
- 01
Built an ingestion pipeline handling PDF, DOCX, Excel and image input, with OCR for scanned pages, normalising all of it into searchable chunks.
- 02
Made project isolation a property of retrieval rather than a filter applied afterwards, so a query cannot reach documents outside its own engagement.
- 03
Implemented retrieval-grounded chat over the indexed corpus using a vector store alongside the relational data, so answers are traceable to source material.
- 04
Moved parsing, embedding and indexing to background workers with a queue, keeping large uploads off the request path.
- 05
Secured the API with token auth and role- and guard-based access control, and shipped it containerised through an automated pipeline.
- React 19
- TypeScript
- FastAPI
- PostgreSQL + pgvector
- LangChain
- Celery + Redis
- Docker
Where it landed
Reviewers query a corpus instead of reading it end to end, and every engagement stays sealed from every other. Heavy documents process in the background while the interface stays usable.
A short conversation with an engineer, not a sales qualification call. If we're the wrong people for it, we'll say so and point you somewhere better.


