Case study
Residential solar pipeline
A data pipeline that profiles California solar installers by comparing what each project was supposed to deliver with what it actually delivered.
Built for solar manufacturers, distributors, and financiers looking for installers worth working with.
- records processed
- 1M+
- installer profiles
- 1,000+
- per 400K records, 4 CPU cores
- ~8 min
- infrastructure, run locally
- $0
The problem
Which solar installers actually perform? Where they work, how big their projects are, and whether their systems produce what was promised is all on public record, split between building permits, utility interconnection filings, and state license data. Nothing connects them.
What we built
- A bot files public records requests every month, then reads the reply emails, extracts the attached PDFs, and runs OCR to turn them into data.
- Web scrapers collect county and city permits, alongside California utility interconnection data and state contractor license records.
- Matching ties each contractor license to its business and its permits, then links permits to utility interconnection applications by location, date, and system size.
- Each installer gets a profile: where they install and are licensed to, projected versus installed kilowatts, projected versus actual production, DC system size, and the panels and equipment they use.
- An AI assistant answers questions by querying the database and writing reports.
The result
Over a million records consolidated into profiles of more than a thousand installers, showing who delivers on their projections and who doesn't.
Infrastructure and cost
- Runs entirely on a local machine: no servers and no cloud bill.
- Processes about 400,000 records in roughly 8 minutes on 4 CPU cores.
- OpenAI embeddings power search across records.
- Built in about three months, part-time.