Job description
Own the caselaw and docket data layer at a seed-stage legal tech startup building an AI litigation platform. Your pipelines turn raw legal data into the structured signal that feeds the product's AI strategy layer.
What you'll do
- Own production pipelines across millions of court records (PACER, NYSCEF, state courts, published opinions)
- Wrangle messy semi-structured inputs (PDFs, scanned filings, XML/HTML) into clean, queryable structures
- Design LLM-assisted extraction workflows to turn unstructured legal text into reliable structured signal
- Feed the judicial intelligence and AI strategy layers alongside the full-stack team
- Handle statistical analysis, SQL optimization, RAG, and semantic search, and operate the AWS data stack within SOC 2 guardrails
Compensation
$180,000 to $220,000 base plus 0.15% to 0.25% equity.
Location
Remote within the U.S., New York City strongly preferred, with at least four hours of daily Eastern overlap and quarterly on-sites in New York.
Requirements
- 4 to 8 years building production data pipelines
- Strong AI-for-data background: RAG, semantic search, LLM-assisted extraction
- Legal data experience: PACER, dockets, court records
- PDF parsing, OCR, and document extraction tooling
- Proficient in Python and SQL for pipeline engineering
- Experience with large, messy datasets, not only clean, pre-structured data
- Bachelor's in computer science, data science, or a related field
- Fully autonomous; able to review code, not just generate it
- U.S.-based; no sponsorship, though visa transfers are accepted
- Vector databases (Pinecone, Voyage AI) a plus