Unlocking unstructured data to power legal AI
One node tree for the entire record, from sub-document down to section, paragraph and line. Schedules, exhibits and attachments keep their own type and page boundaries. Reading end-to-end resolves what a page-by-page parser guesses at: whether the table on page 47 started on page 34.

LexSelect identifies each element for what it is, including struck text, signature blocks, seals, footnotes, page numbers, handwriting, images, charts, tables and passage language. Non-body elements remain distinct, and struck language is not treated as operative text.

Citations follow the document’s own reference system, not only page coordinates. Cite subsection 12.3(a) of Schedule A, paragraph 47 of Exhibit D within an affidavit, or deposition transcript p. 4, ll. 20-21.

Send anything from a two-page letter to a 1,800-page minute book to the API. LexSelect returns a structured node tree with every element classified, nested and linked to its source. Build on that and your outputs stay consistent, stay traceable, and cost less to run.

Clean text
Reading-order text from any layout, including multi-column pages.

Structure & hierarchy
Headings, clauses, numbering, and nesting: the document tree, not a text dump.

Tables
Rows, headers, and merged cells, stitched across page breaks.

Handwriting
Handwritten annotations, notes, and form entries on scanned documents.

Signatures & stamps
Signature and stamp detection with page locations.

Document types
Automatic classification: the engine detects what each document is on its own.

Document relationships
Exhibits linked to filings; sub-documents inside compiled bundles detected and separated.

Source citations
Exact page and position for every extracted element.


Affidavits
Sworn statements with their exhibits attached. Numbered paragraphs are preserved and each exhibit is detected as its own sub-document, linked to its parent.
Annotated elements

Court filings
Complaints, motions, orders, petitions: anything filed with a court, typed or handwritten. Form fields come back as key-value pairs and docket stamps stay out of the text.
Annotated elements

Transcripts
Depositions, hearings, and examinations. Every line is numbered, every question paired with its answer, ready to cite by page and line.
Annotated elements

Contracts
Agreements of any kind, from employment to commercial. Clause hierarchy survives the parse: 12.3(a) stays 12.3(a), and schedules are detected as sub-documents.
Annotated elements

Legislation
Bills, statutes, and regulatory issues: deeply nested numbering, definitions, and cross-references, returned as the hierarchy the text was written in.
Annotated elements




More
Notices, wills, corporate records, court forms, and everything in between. The classification vocabulary is open: the engine detects what each document is on its own.





LexStudio / API
Developer portal
Generate an API key, read the docs, and make your first parse in minutes.
Upload your own documents in the Playground and inspect every extracted class and structure visually.
Track usage against your plan and review logs, so you can see exactly what was processed and what it cost.


LexChat
AI legal assistant
Copy clean, perfectly formatted text from any document, even scanned or complex PDFs, and generate citations for any passage.
Ask questions across all your uploaded documents and get accurate, cited answers that link back to the original source.
Use it inside Microsoft Word with our add-in: review documents, run LexChat, and insert cited text without leaving your draft.
LexEnterprise
Enterprise pipelines
Custom intake, extraction, and review workflows, configured for your process.
Delivered through the same products your team already uses, with structured data flowing into your systems.

If you're building a product, start with LexStudio: generate an API key and parse your first document in the Playground. If you work with documents day to day, start with LexChat: sign up free and put it to work on a real matter. Teams with volume or custom workflow needs can contact us directly.
General-purpose tools extract text and tables, but the hierarchy, the relationships, and the legal context are lost. LexSelect returns a structured tree built for legal documents: exhibits detected as sub-documents and linked to their parent, tables tracked as one table across page breaks, transcripts paired question to answer, and every document classified automatically. See the full comparison on the API page.
Open LexStudio, create an account, and generate a key. Standard terms are accepted on activation, so there's no contract to negotiate before you start building.
PDF, Word (DOC, DOCX), RTF, ODT, HTML, and email files (EML, MSG), up to 250 MB per file, scanned or born digital.
The parse response is JSON: a structured tree per page, plus optional slices like plain text, extracted tables, and key-value pairs. A separate render call returns Markdown or HTML. Every element carries its exact position on the page.
LexSelect is SOC 2 Type II audited, with encryption in transit and at rest and strict data isolation. See the privacy policy and terms for the specifics.
Contact us and we'll scope it together, from high-volume archives to document intelligence workflows configured for your process.

SOC 2 Type 2
audited
Every document you process on LexSelect is protected by independently audited security controls designed to keep sensitive legal data secure.










