Back to AI Labs

AI LabRetrieval

The assistant on this site.

The field above and the ghost in the corner are one assistant over this portfolio. Code decides and the model only writes: a separate judge settles whether a question can be answered before a word is generated, and the answers that must never be wrong come straight from a table.

Role
Built it end to end
Retrieval
Postgres, vector and keyword
Judge
Jev by TypeSafe
Writer
Claude Haiku 4.5
01

The problem

A portfolio answers the questions its author anticipated. Anything else means reading every page. The obvious fix, putting a chatbot on top, usually makes it worse: a model with a friendly prompt will confidently invent a job title or a date. For a site whose whole purpose is an accurate record of what I have done, a confident wrong answer is worse than no answer.

What is Antonis' current job?

A chatbot with a friendly prompt

Antonis is a Senior Product Manager at Suitsupply, where he has led the eCommerce roadmap since 2021.

Neither the title nor the date is anywhere on this site.

This assistant

Antonis is currently working as Junior Business Analyst at Suitsupply. Here's the full picture:

  • Junior Business Analyst at Suitsupply, Amsterdam (Jun 2026 – Present)
  • Customer Service Specialist at Suitsupply, Amsterdam (Oct 2025 – May 2026)

From the experience table. The model is never called.

02

What I built

Retrieval over the site's own content, with a hard line between what the model may phrase and what it may never decide.

RoleCompanyDates
Junior Business AnalystSuitsupplyJun 2026 –
Customer Service SpecialistSuitsupplyOct 2025 – May 2026
Marketing OfficerUnicafSep 2024 – Mar 2025

Facts without the model

Job titles, employers and dates come from a table seeded from the repository. No model call, so no invented employer or year.

Vector
1. AI Labs2. Omniscope3. Home
Keyword
1. Omniscope2. Cloudflare3. Atlas
Merged
1. Omniscope2. AI Labs3. Cloudflare

Two searches, one ranking

Every page is indexed twice, as vectors for meaning and as words for names. Both searches run in one database call and their rankings are merged, so a question about Omniscope finds the page that names it.

Answerable0.95
Refusal line 0.30answers

A judge before the writer

One request to Jev, a model that returns probabilities instead of text, settles the route, the reply language, whether the evidence answers the question and which passages to keep.

I don't have that specific information available in my knowledge base.

A refusal that is fixed

Below the answerable line the reply is a fixed sentence and a contact form. The model is never asked, so it cannot improvise one.

enes

He speaks English fluently, Spanish at B1 and Greek natively.

Your language

Routing, refusals and the job list work the same in English and Spanish, and the model replies in the language you wrote in.

38

of 40 passed

0 answerable questions refused

Measured, then logged

Forty questions in two languages, traps and injection attempts included, run against every change. Each live reply is logged with its route and scores, personal data masked.

03

How a question is answered

Every reply reports what was decided for it. This is a real one from the eval. Ask the assistant at the top of the page and it becomes yours.

Ask the assistant at the top of the page to see your own question here.

What did he do at Unicaf and when?

Antonis was Marketing Officer at Unicaf in Cyprus from September 2024 to March 2025. He executed international marketing campaigns across multiple African markets, optimised landing pages with UTM tracking, delivered email marketing and CRM workflows using Brevo and Elastic Email, and built the on-ground marketing wiki in Notion. His work included the Unicaf Email Campaigns promoting scholarship opportunities across 5+ markets, retention and engagement flows with 3+ automated lifecycle journeys, and a comprehensive Notion wiki that achieved 50% faster approvals and 100% team adoption across 30+ processes.

  1. Retrieve

    378 ms

    12

    candidates · vector and keyword

    The question is embedded and matched against the page chunks twice, by meaning and by words, in one call. Twelve candidates come back.

  2. Judge

    179 ms

    RouteJobs and dates99%

    Answerable0.95

    Answers · Kept 8 of 12

    Jev reads the question, the profile and the candidates, and returns the route, the reply language, the probability the evidence answers it, and a relevance for each passage.

  3. Reply

    1,169 ms

    Written by Claude Haiku 4.5

    Code picks the reply. Off topic or below the line: a fixed sentence. A job list: straight from the table. Anything else: the model writes from the passages that were kept.

  4. Log

    One row, personal data masked

    One row per reply: the question with emails and phone numbers masked, the route, the scores and the timings. Nothing that identifies the visitor.

Where the time goesFirst word 1,169 ms · Total 3,646 ms
Retrieve 378 msJudge 179 msReply 3,089 ms

Passages the judge kept

  • Home0.92
  • Retention engagement flows0.91
  • Unicaf wiki0.89
  • Unicaf campaigns0.88
  • Unicaf campaigns0.85
  • Ai innovation0.83
04

My role

I designed and built all of it: the index, the retrieval, the judge, the API route, both chat surfaces and the eval.

  • 01

    Where the model may not decide

    Decided which answers come from tables and which replies are fixed, and built those paths before the model's.

  • 02

    The judge

    Wrote the Jev questions for routing, answerability and passage relevance, and set the thresholds against the eval.

  • 03

    Retrieval and the index

    Hybrid search in Postgres, and a weekly reindex that swaps the new index in with one transaction.

  • 04

    The eval

    Forty questions in English and Spanish, traps and injection attempts included. The old pipeline passed 17; this one passes 38.

  • 05

    Both surfaces

    The corner assistant and this page share one conversation hook, one renderer and one trace.

05

What changed

38/40

Eval questions passed, up from 17

Measured on the same forty questions before and after, against the live site and then the new pipeline.

Before
After
Rules that only read English
One judge, in either language
A model free to decide it knew
A threshold decides, the model writes
Sixty chunks for every question
Only the passages the judge kept
No way to tell if a change helped
An eval run against every change

0

Answerable questions refused, down from 11

24

Pages indexed

~180 ms

For the judge to decide

06

What I would do differently

The index is still rebuilt in full every week, which is simple and safe but wasteful: change tracking would re-embed only the pages that changed. The conversation history still comes from the browser, so a visitor can rewrite what the assistant told them earlier; it belongs on the server. And the judge's thresholds were set on forty questions, which is enough to catch a regression and not enough to call them tuned.

07

Built with

  • TypeScript
  • Next.js
  • Supabase
  • PostgreSQL
  • pgvector
  • Claude
  • Jev by TypeSafe
  • OpenAI embeddings
  • Firecrawl
  • Upstash Redis
  • Vercel