S4E1, season 4, monsoon
I asked my rejections what they had in common
ROLEFIT
Retrieval-augmented generation appears in 7 of the 9 rejections and 1 of the 4 that replied. Vector databases, usually Pinecone or pgvector, appear in 6. Both are absent from every role that reached a screen.
Why
I was applying to a lot of AI engineering roles and getting a lot of skill-mismatch rejections without ever being told which skill. The job descriptions were sitting in my browser history the whole time, so I put them in a database and asked them directly.
How it's built
- Hybrid retrieval. Dense embeddings match meaning, Postgres full-text matches rare exact words like tool names, and the two rankings fuse with Reciprocal Rank Fusion. The whole thing is one SQL function.
- Corrective retrieval. A grader reads the chunks before anything is generated. If they can't answer the question, it rewrites the query and retrieves again, up to twice.
- An admit-gap branch. Out of attempts, it says the corpus doesn't cover the question. That branch is the point.
- Embeddings with no embeddings vendor: gte-small, running inside a Supabase Edge Function.
- No database connections from serverless. Everything goes through PostgREST and security-definer functions, rate limited in Postgres.
The eval
The eval suite gates on deterministic citation validity. That is how I found the harness bug that was feeding the LLM judge the wrong input and scoring correct answers 0 out of 4.
Python, LangGraph, Supabase, Postgres, gte-small, Groq
Next episode, S4E2
Hiring before deciding