Automation
A pipeline that knows when it's guessing.
A sourcing and matching pipeline, and its test case was a job search: it finds postings, scores them and prepares applications. The same problem shows up in any messy inflow, like leads, vendors, listings or candidates.

- The problem
- Over twenty sources were closed, and the free ones that worked gave only 2 to 6 relevant matches a day.
- What I built
- A pipeline that reads 8,000 to 9,600 postings a run, scores each one and explains the score.
- The result
- About 90 matches a day, up from 6, all waiting in a review queue.
How I built it
The problem
Most job boards are walls.
The constraint was strict: everything free. No paid APIs and no paid model calls. Under that rule, most of the obvious sources closed straight away.
I probed over twenty boards. Some needed keys, some returned 403 behind Cloudflare, and some just returned 404. Everything that worked, added together, produced between two and six relevant roles a day.
What actually worked
One logged-out endpoint.
The source that changed it was LinkedIn's public guest search, the same pages anyone sees without logging in. It needs no session scraping and no login. It was also the only source carrying the roles that actually fit, and it took the queue from about 6 a day to about 90.
A daily run now reads roughly 8,000 to 9,600 postings and keeps the top 120.
What I built
A score is only useful if it explains itself.
Volume brought a new problem. A teaser card from a search page is not a job description, and fetching every full description is slow, so some matches get scored on the title alone.
Those are labelled title-only match on the dashboard. A match and a lead are different things, and the dashboard should say which one you are looking at.
Scored on what the role does
A flat reject list once killed 'Sales Operations Analyst' because of the word sales. About 149 real postings were lost that way.
Only listed locations get in
Block-lists leaked twice. 'Remote, US' is still not remote for someone in India, so locations now have to be on an allow-list.
Six archetypes, 103 distinct letters
One template produced 42 near-identical letters. Each archetype opens with a different piece of real evidence.
It will not overclaim
Experience is only claimed when the evidence repeats, so three incidental mentions can't make one up.
What is not done yet
It never presses submit.
Some of what is missing here is missing on purpose.
- There is no auto-apply, on purpose. It is a review queue, and a person presses submit.
- The location filter trusts employers' own 'remote' flag, which is often wrong, so some results still show a city.
- It runs on one machine. There is no hosted version.