Back to the series

S4E1, season 4, monsoon

I asked my rejections what they had in common

INT. PARI'S ROOM - NIGHT
July 2026
Personal project, live

ROLEFIT

Retrieval-augmented generation appears in 7 of the 9 rejections and 1 of the 4 that replied. Vector databases, usually Pinecone or pgvector, appear in 6. Both are absent from every role that reached a screen.

Why

I was applying to a lot of AI engineering roles and getting a lot of skill-mismatch rejections without ever being told which skill. The job descriptions were sitting in my browser history the whole time, so I put them in a database and asked them directly.

How it's built

  • Hybrid retrieval. Dense embeddings match meaning, Postgres full-text matches rare exact words like tool names, and the two rankings fuse with Reciprocal Rank Fusion. The whole thing is one SQL function.
  • Corrective retrieval. A grader reads the chunks before anything is generated. If they can't answer the question, it rewrites the query and retrieves again, up to twice.
  • An admit-gap branch. Out of attempts, it says the corpus doesn't cover the question. That branch is the point.
  • Embeddings with no embeddings vendor: gte-small, running inside a Supabase Edge Function.
  • No database connections from serverless. Everything goes through PostgREST and security-definer functions, rate limited in Postgres.

The eval

The eval suite gates on deterministic citation validity. That is how I found the harness bug that was feeding the LLM judge the wrong input and scoring correct answers 0 out of 4.

Python, LangGraph, Supabase, Postgres, gte-small, Groq

Next episode, S4E2

Hiring before deciding