Engineering service
Hybrid RAG & retrieval engineering
Vector + relational retrieval, semantic chunking, and re-ranking so LLM outputs stay grounded when catalogues and documents get large.
- Stores
- Vector + Postgres / IndexedDB
- Focus
- Chunking · ranking · evals
- Typical slice
- 2–6 weeks per pipeline
In depth
Retrieval is where most “hallucination” problems actually start: wrong chunks, stale serialisation, or ranking that ignores query intent. I build hybrid retrieval that combines vector search with relational filters and metadata you already trust, then tie changes to evals so you know when quality moves.
At Parker AI this meant Qdrant plus Supabase Postgres, chunking tuned for long-form social and ads content, and query-aware re-ranking. At Magic, recall also spans assistant playbooks and tickets — including patterns that keep sensitive context local when possible.
What you get
Schema design for vectors, metadata, and relational joins
Chunking and embedding conventions for your content shapes
Ranking and re-ranking strategy tied to real queries
Eval harness and tracing hooks for retrieval-led regressions
Rollout plan: backfill, indexing throughput, and monitoring
How we work
Baseline the failures
When does the model go generic? Trace it to retrieval, not “the model.”
Model the data
What belongs in vectors vs facts you filter in SQL?
Iterate chunk + rank
Ship changes behind evals; compare before/after on held-out queries.
Operational hygiene
PII boundaries, retention, and consistent serialisation for agents.
Examples & past work
Outcomes you can expect
- Design hybrid stores (e.g. Qdrant + Postgres) with query-aware re-ranking and consistent serialisation.
- Tune embedding and chunking pipelines for long-form social, ads, and document content.
- Pair retrieval changes with evals and tracing so quality regressions are measurable.
Questions, answered
Do I need both a vector DB and Postgres?
Often yes for real products: vectors for similarity, Postgres for authoritative filters, tenancy, and fields you do not want embedded.
How do you measure retrieval quality?
Golden questions, side-by-side ranking checks, and downstream task success (e.g. correct citations or tool args) — wired into traces so regressions show up in CI or review dashboards.
Can you work with our existing embeddings vendor?
Yes. The important part is consistent preprocessing and evaluation — not a specific brand of embeddings API.
Book a free 30-minute discovery call to investigate your work and needs
I will map constraints, risks, and a practical first milestone — whether that is agents, retrieval, ingestion, extensions, or full-stack SaaS delivery.
Other services
Multi-agent workflows with planning, memory, typed tool calls, and human checkpoints — from creative-strategy platforms to assistant co-pilots.
Schedulers, retries, and normalised pipelines from ads APIs, social platforms, and internal services — built for scale and observability.
Chrome extensions and sidepanel experiences with in-browser RAG and real actions via tool gateways — not another detached chat tab.