[{"data":1,"prerenderedAt":542},["ShallowReactive",2],{"blog-post-\u002Fblog\u002Flet-the-llm-read":3,"blog-posts":539},{"id":4,"title":5,"audio":6,"audioDuration":7,"body":8,"cover":521,"date":522,"description":523,"draft":524,"extension":525,"meta":526,"narrationHash":527,"navigation":528,"path":529,"readingTime":530,"seo":531,"stem":532,"tags":533,"__hash__":538},"blog\u002Fblog\u002Flet-the-llm-read.md","Let the LLM read. Let math decide.","\u002Faudio\u002Flet-the-llm-read.mp3",1142,{"type":9,"value":10,"toc":506},"minimark",[11,19,22,25,28,50,53,60,63,66,71,74,77,85,88,95,99,102,119,122,125,128,154,157,161,164,167,170,184,187,204,207,211,214,229,232,236,239,242,245,294,297,300,304,307,310,321,328,332,335,338,341,358,362,365,368,371,374,377,380,384,387,413,416,420,458,462,500,503],[12,13,14,18],"p",{},[15,16,17],"strong",{},"Profilon"," matched CVs with job ads, for people looking for their next role and for companies looking for the right people.",[12,20,21],{},"We shut it down. There wasn’t enough interest, and companies had already started building similar tools of their own. None of them came close to our matching engine, but we decided that SaaS in this field wasn’t the right call, closed the project and moved on to other projects and ideas.",[12,23,24],{},"The matching engine is still worth writing about. The approach works for anyone building ranking or matching with AI, and it saved us a lot of money.",[12,26,27],{},"When we started, the obvious design was to hand an LLM a CV and a job ad and ask how good the match is, from 0 to 100. It looks great in a demo. In production it has three problems.",[29,30,31,38,44],"ul",{},[32,33,34,37],"li",{},[15,35,36],{},"It gets expensive fast."," Matching is about pairs. Every new CV has to be checked against every open job, and every new job against every CV.",[32,39,40,43],{},[15,41,42],{},"It isn’t deterministic."," The same pair can get 72 today and 64 tomorrow. Good luck explaining that to a recruiter.",[32,45,46,49],{},[15,47,48],{},"It can’t explain itself, and it can be talked into things."," CVs are untrusted input, and hiding instructions for AI screeners in white text is already a known trick.",[12,51,52],{},"So we ended up with one rule.",[54,55,57],"blog-rule",{"label":56},"The rule",[12,58,59],{},"The LLM is allowed to read. It is never allowed to score a match.",[12,61,62],{},"The LLM turns each CV and each job ad into structured data, once. Everything after that runs on regular CPUs, using small models and plain old techniques: finding candidates, scoring them and deciding what counts as a match.",[64,65],"llm-pipeline-figure",{},[67,68,70],"h2",{"id":69},"pairs-are-the-expensive-part","Pairs are the expensive part",[12,72,73],{},"The number of documents grows in a straight line. The number of pairs grows with the product.",[12,75,76],{},"Say you have 1,000 CVs and 10,000 job ads. That’s 11,000 documents, but 10 million possible pairs. If an LLM scores every pair, that’s 10 million calls. A smarter setup only lets the LLM judge each CV’s top 50 jobs. That’s still 50,000 calls, and you pay for them again every time a CV is updated, a job ad changes or someone tweaks the prompt.",[12,78,79,80,84],{},"In our setup the LLM reads each document once. That’s 11,000 calls, and then never again for that document. Changing how we score ",[81,82,83],"highlight",{},"costs zero LLM calls",", because scoring never used one.",[86,87],"llm-calls-figure",{},[12,89,90,91,94],{},"So the rule for cost is simple. Anything you pay for per call should happen ",[81,92,93],{},"once per document, never once per pair",".",[67,96,98],{"id":97},"_1-read-the-only-llm-step","1. Read: the only LLM step",[12,100,101],{},"We used Gemini 2.5 Flash, one of Google’s smaller and cheaper models. It gets a strict JSON schema and a low temperature (0.1), and it gives back things like:",[29,103,104,107,110,113,116],{},[32,105,106],{},"skills with a level from 1 to 5, per role, so we know when each skill was last used",[32,108,109],{},"job titles and years of experience",[32,111,112],{},"seniority",[32,114,115],{},"languages with CEFR levels, like B2 or C1",[32,117,118],{},"industries, certifications and education",[12,120,121],{},"For job ads it also separates required skills from nice-to-haves, and it marks the one to three skills the role can’t do without.",[12,123,124],{},"It translates too. A lot of CVs and job ads in Sweden are written in Swedish. Everything comes out in English, so the rest of the pipeline can get by with a small English model.",[12,126,127],{},"A few more things kept the cost down:",[29,129,130,136,142,148],{},[32,131,132,135],{},[15,133,134],{},"Batch API for job ads."," Job ads don’t need an answer in two seconds, so they go through Google’s batch API at about half the price.",[32,137,138,141],{},[15,139,140],{},"PDFs go straight in."," Gemini can read a PDF directly. By our estimate that used 2 to 5 times fewer tokens than sending the extracted text, and it copes with messy layouts and scanned CVs without a separate OCR step.",[32,143,144,147],{},[15,145,146],{},"Only parse what changed."," Job boards get scraped every day. An ad only goes back to the LLM if its text actually changed.",[32,149,150,153],{},[15,151,152],{},"Pick the right prompt first."," Before the LLM sees anything, a small multilingual model on the CPU sorts the text into one of ten domains, like technology, healthcare or logistics. It compares the text with example sentences in English and Swedish, so it needed no training. The domain decides which extraction rules go into the prompt. A nurse’s CV and a backend developer’s CV need very different rules, and specific rules are what let a cheap model do a good job.",[12,155,156],{},"The result is stored in the database. From here on, the LLM is done.",[67,158,160],{"id":159},"_2-turn-words-into-vectors-on-a-cpu","2. Turn words into vectors, on a CPU",[12,162,163],{},"An embedding model turns a piece of text into a list of numbers, a vector. Texts that mean the same thing end up close to each other, even when they share no words. “PostgreSQL” and “Postgres” land right next to each other.",[12,165,166],{},"We used BAAI’s bge-base-en-v1.5 through the sentence-transformers library. It’s a BERT model with about 110 million parameters. BERT came out in 2018, which is ancient in AI years, and it’s still great at this job. It turns any text into 768 numbers and runs fine on an ordinary CPU.",[12,168,169],{},"We embedded two kinds of things:",[29,171,172,178],{},[32,173,174,177],{},[15,175,176],{},"Every skill name on its own."," Our skills table grew past 90,000 unique names. People write React, React.js and ReactJS, and exact string matching misses a lot of real matches.",[32,179,180,183],{},[15,181,182],{},"A short summary of every CV and every job ad."," Title, summary, recent roles and skills for a CV. Title, role, description and skills for a job.",[12,185,186],{},"And we made it cheap to run:",[29,188,189,192,195,198,201],{},[32,190,191],{},"Our workers were ordinary ARM servers. No GPUs.",[32,193,194],{},"The models run on ONNX Runtime instead of PyTorch. It has optimized kernels for those ARM CPUs, and the outputs were numerically identical.",[32,196,197],{},"We installed the CPU-only build of PyTorch, which cut about 2 GB of CUDA packages we’d never use from every Docker image.",[32,199,200],{},"The models are baked into the Docker image, so a new worker starts without downloading anything.",[32,202,203],{},"The vectors live in Postgres with pgvector, stored as half-precision floats. That halved the size of the vectors and the index with barely any effect on search quality.",[12,205,206],{},"There’s no separate vector database. Postgres handled keyword search, vector search and full-text search, and that’s one less system to run, back up and pay for.",[67,208,210],{"id":209},"_3-find-candidates-with-keywords-and-vectors","3. Find candidates with keywords and vectors",[12,212,213],{},"Before scoring anything we narrow it down with two cheap searches in the same database:",[215,216,217,223],"ol",{},[32,218,219,222],{},[15,220,221],{},"Keyword overlap."," Plain SQL: which jobs share at least one skill name with this CV?",[32,224,225,228],{},[15,226,227],{},"Vector search."," Which 100 jobs are closest in meaning to this CV? pgvector’s HNSW index keeps that fast.",[12,230,231],{},"Keyword search catches the obvious matches. Vector search catches the ones that use different words, like “Backend Developer” and “Software Engineer, APIs”. We merge the two lists, drop duplicates and only score what’s left.",[67,233,235],{"id":234},"_4-score-it-the-old-fashioned-way","4. Score it the old-fashioned way",[12,237,238],{},"This is where a lot of teams would reach for the LLM again. We didn’t. The base score is a weighted sum of eight signals. Skills are 66% of it, experience and seniority 8% each, language, industry, education and domain 4% each, and certifications 2%.",[12,240,241],{},"Skills carry most of the weight, so they get most of the logic.",[243,244],"llm-skill-match-figure",{},[29,246,247,253,259,265,271,277,283,288],{},[32,248,249,252],{},[15,250,251],{},"Everything against everything."," Each skill in the job ad is compared with each skill in the CV using cosine similarity. It’s one small matrix multiplication.",[32,254,255,258],{},[15,256,257],{},"One CV skill per requirement."," The best pairs are assigned first and each CV skill can only be used once, so one strong skill can’t cover five requirements.",[32,260,261,264],{},[15,262,263],{},"0.82 is a match, 0.90 is a strong match."," We started at 0.75 and got too many false positives. We went up to 0.86, then back down to 0.82, where it stayed.",[32,266,267,270],{},[15,268,269],{},"Check the text as well."," If the CV doesn’t list Terraform as a skill but a role says “moved our infrastructure to Terraform”, Postgres full-text search finds it and it counts as a partial match. Full-text search has been built into Postgres since 2008.",[32,272,273,276],{},[15,274,275],{},"Small adjustments."," A bonus when the candidate’s level meets the required level. Extra weight for certified skills. Less weight the longer ago a skill was used, but never below 30%. A bonus for years of use on a log scale, so ten years isn’t worth ten times one year. And the one to three skills the LLM marked as essential count 1.5 times as much.",[32,278,279,282],{},[15,280,281],{},"Discount the fluff."," “Communication”, “teamwork” and “Microsoft Office” count for a quarter of a normal skill. We built that list from production data, from the skills behind the most false positives. Everyone lists them, so they say almost nothing about fit.",[32,284,285],{},[15,286,287],{},"Required skills count 85%, nice-to-haves 15%.",[32,289,290,293],{},[15,291,292],{},"A sanity check on the whole profile."," We average the CV’s skill vectors and the job’s skill vectors and compare the two averages. If they point in different directions, the score drops by up to 15%. A CV that shares three keywords with a job but is really about something else gets caught here.",[12,295,296],{},"The other signals are simple rules. Years of experience against the requirement. Distance between seniority levels. CEFR levels for languages. Same, related or different domain.",[12,298,299],{},"None of this is new. Cosine similarity, full-text search, weighted sums and log scales have been around for decades. That was on purpose. Every number can be explained, tested and tuned.",[67,301,303],{"id":302},"_5-a-second-opinion-from-a-small-model","5. A second opinion from a small model",[12,305,306],{},"Embeddings compare things one at a time. They never read the job and the CV together. For that we used a cross-encoder called ms-marco-MiniLM-L12-v2. It’s a distilled BERT-style model with about 33 million parameters, trained on Microsoft’s MS MARCO search dataset. It reads the job and a short profile summary as one pair, the way a search engine ranks results, and returns a relevance score.",[12,308,309],{},"It runs on the CPU, in one batched call, and only on the top 50 candidates by base score. We treated it as a second opinion:",[29,311,312,315,318],{},[32,313,314],{},"It’s worth 5% of the final score.",[32,316,317],{},"It can veto. Below 30 out of 100, the match is dropped.",[32,319,320],{},"It can’t rescue a weak match. When the skills don’t line up, its score is capped at twice the skill score, so an AI score of 100 on a skill score of 20 counts as 40.",[12,322,323,324,327],{},"So yes, there’s a model judging the top 50 after all. It’s just ",[81,325,326],{},"a 33-million-parameter model on our own CPU"," instead of a paid API call.",[67,329,331],{"id":330},"_6-the-final-score","6. The final score",[12,333,334],{},"The final score blends everything: 35% base score, 25% skill quality (how close the skill matches are and how specific the skills are), 15% job title similarity, 8% experience, 6% seniority, 5% AI, 4% industry and 2% language.",[336,337],"llm-score-weights-figure",{},[12,339,340],{},"Then a few rules on top:",[29,342,343,346,349,352,355],{},[32,344,345],{},"If the job titles are far apart, up to 12 points come off.",[32,347,348],{},"If the domains don’t match, the match is blocked. We added that after an AI engineer with some healthcare in their background kept getting matched with nurse and psychologist jobs. The domain check blocked 9 of the 10 bad matches we tested it on.",[32,350,351],{},"If a required language is missing, the match is hidden.",[32,353,354],{},"Only matches that score 65 or higher are shown.",[32,356,357],{},"When the base score and the AI agree, confidence goes up. When they disagree it goes down, and the match can drop a tier.",[67,359,361],{"id":360},"_7-tune-it-with-votes-check-it-with-agents","7. Tune it with votes, check it with agents",[12,363,364],{},"Weights like these are never right on the first try, so we let the users tell us where they were wrong.",[12,366,367],{},"Every match had an up vote and a down vote. A down vote needed a short reason, and a down-voted match disappeared from the list. We gathered the votes and used them to retune the weights and thresholds automatically.",[12,369,370],{},"Tuning on votes alone is risky, though. A change that fixes one kind of bad match can quietly break another. So after a retune we let LLM agents look at the results and answer one question: did this make the matches better, or worse?",[12,372,373],{},"That’s the one place an LLM gets to judge anything, and it sits outside the live pipeline. It reviews changes to the scoring. It never scores a match.",[12,375,376],{},"Automatic tuning also needs an undo button. Every match stores the version of the scoring config that produced it. The config itself is versioned, and a revert creates a new version from an old one, so nothing gets lost. If a retune made things worse, we could see exactly which matches it had scored, trace any score back to the weights behind it, and roll back.",[378,379],"llm-tuning-loop-figure",{},[67,381,383],{"id":382},"basically-a-small-search-engine","Basically a small search engine",[12,385,386],{},"Looking back, we built a small search engine where the query is a whole document. A CV searches for jobs, and a job ad searches for CVs. The big web search engines have run on the same ideas for years:",[29,388,389,395,401,407],{},[32,390,391,394],{},[15,392,393],{},"Read once, answer many times."," Google doesn’t read the web when you search. It crawls pages ahead of time, stores what it learned in an index and ranks from that index when you search. Pages that change get crawled again. Our LLM step is the crawl: it reads each document once, and again only if the text changes.",[32,396,397,400],{},[15,398,399],{},"Keywords plus meaning."," Search engines combine exact keyword matching with semantic matching, so they can find pages that use different words than your query. That’s our keyword overlap plus vector search. The keyword side is the classic inverted index idea: from each skill to every job that asks for it.",[32,402,403,406],{},[15,404,405],{},"Cheap first, expensive last."," A fast first pass cuts a huge pile down to a short list, and heavier models only rank the top of it. Google started using BERT in Search in 2019. We used a small BERT-style cross-encoder on our top 50, and it was trained on MS MARCO, a dataset Microsoft built from real Bing searches.",[32,408,409,412],{},[15,410,411],{},"Test every change."," Google has human quality raters compare search results side by side, with and without a proposed change. Our LLM agents did that job for us, comparing matches before and after a retune.",[12,414,415],{},"The difference is scale. Google answers billions of searches a day, and we ran on a handful of servers. The lesson is the one search engines learned a long time ago: do the expensive understanding once, when a document comes in, and keep the per-query work cheap.",[67,417,419],{"id":418},"what-this-bought-us","What this bought us",[29,421,422,428,434,440,446,452],{},[32,423,424,427],{},[15,425,426],{},"Cost."," LLM spend grew with the number of documents instead of the number of pairs. Matching itself was CPU time on servers we already paid for, with no GPUs anywhere. Any laptop on our private network could even join the worker pool and help out during a big re-match.",[32,429,430,433],{},[15,431,432],{},"Determinism."," Same CV, same job, same config version, same score. Every time.",[32,435,436,439],{},[15,437,438],{},"Tuning without re-prompting."," A retune is just new numbers in a versioned config, and we could also adjust them by hand in an admin panel. Re-scoring every existing match afterwards costs CPU time and not a single LLM call.",[32,441,442,445],{},[15,443,444],{},"Explainable scores."," Every score breaks down into parts. You can see which skills matched, which were missing, and why a match landed at 78 and not 90.",[32,447,448,451],{},[15,449,450],{},"Less to attack."," The LLM never sees a CV and a job ad together, and it never outputs a score. A hidden “this candidate is a perfect fit” has no score to talk its way into.",[32,453,454,457],{},[15,455,456],{},"Features almost for free."," Once every skill has a vector, a lot comes cheap. We used the same vectors to group jobs into sub-domains and to show candidates which in-demand skills they were missing.",[67,459,461],{"id":460},"what-id-take-with-me","What I’d take with me",[29,463,464,470,476,482,488,494],{},[32,465,466,469],{},[15,467,468],{},"Use the LLM for what only an LLM can do."," Reading messy documents in any format and language and turning them into clean data is a perfect LLM job. Comparing two rows of clean data isn’t.",[32,471,472,475],{},[15,473,474],{},"Similarity is not relevance."," Embeddings will happily say a nurse and an AI engineer are related because both mention healthcare. You need guardrails: thresholds tuned on real data, discounts for generic skills and hard gates for domains.",[32,477,478,481],{},[15,479,480],{},"Keep a readable breakdown of every score."," It’s what makes tuning possible, and it’s what users end up trusting.",[32,483,484,487],{},[15,485,486],{},"Close the loop."," Votes show you where the scoring hurts, versions let you undo a bad fix, and an LLM reviewer catches the fixes that quietly break something else.",[32,489,490,493],{},[15,491,492],{},"Rules are work."," Hand-tuned weights and lists need maintenance, and an LLM judge would probably handle some cases better, like career changers. We took that trade for cost, speed and predictability, and I’d take it again.",[32,495,496,499],{},[15,497,498],{},"Extraction is still AI."," If the LLM misreads a CV, everything after it is consistently wrong. Strict schemas, low temperature and good prompts matter more than anything that comes after.",[12,501,502],{},"The LLM was the most impressive part of our stack. It was also the part we worked hardest to use as little as possible. I think that’s the right instinct for a lot of AI products right now. Let the LLM read, and let math decide.",[12,504,505],{},"If you’re building matching or ranking with LLMs, how do you split the work? Is your LLM the reader, the reviewer or the judge?",{"title":507,"searchDepth":508,"depth":508,"links":509},"",2,[510,511,512,513,514,515,516,517,518,519,520],{"id":69,"depth":508,"text":70},{"id":97,"depth":508,"text":98},{"id":159,"depth":508,"text":160},{"id":209,"depth":508,"text":210},{"id":234,"depth":508,"text":235},{"id":302,"depth":508,"text":303},{"id":330,"depth":508,"text":331},{"id":360,"depth":508,"text":361},{"id":382,"depth":508,"text":383},{"id":418,"depth":508,"text":419},{"id":460,"depth":508,"text":461},"\u002Fimages\u002Fblog\u002Flet-the-llm-read\u002Fcover.png","2026-09-30","How we matched CVs to job ads at Profilon using sentence transformers on plain CPUs, and why the LLM never got to set the score.",false,"md",{},"b71022c0059ba485",true,"\u002Fblog\u002Flet-the-llm-read",13,{"title":5,"description":523},"blog\u002Flet-the-llm-read",[534,535,536,537],"AI","LLM","Search","Postgres","ABmDJ3w4zalQc2QTHPK1lIFagENn3Ah8vKBbeEOdNxc",[540],{"path":529,"title":5,"description":523,"date":522,"cover":521,"tags":541,"readingTime":530,"audio":6},[534,535,536,537],1791521608449]