RAG assistants
Chat and search that answer from your own pages, PDFs and databases, cite where the answer came from and hand over to a human when they should.
AI implementation
We build retrieval-augmented (RAG) assistants, AI features and automations that are grounded in your own content, measured against real questions and cheap enough to run every day.
RAGEmbeddingsHybrid searchGeminiOpenAIClaudePython
Overview
Most AI projects that disappoint do so for boring reasons: the assistant was never given the right content, nobody defined what a good answer looks like, and running costs were an afterthought. We treat AI as an engineering problem first — retrieval, evaluation, access control and cost — and a model choice second.
A typical engagement starts small. We take a slice of your content and a list of the questions people genuinely ask, build a working assistant, and measure how well it answers. You see real results on your own material before committing to a wider rollout.
When it goes live, it stays honest: answers come from your content, sources can be shown, gaps are logged so they can be filled, and the model behind it can be swapped as prices and capabilities change. The assistant on this site is itself an example of the approach.
What we deliver
From a single assistant to AI woven through your product — each piece built to be measured, not just demoed.
Chat and search that answer from your own pages, PDFs and databases, cite where the answer came from and hand over to a human when they should.
Summaries, smart search, classification, drafting and extraction added to your web or mobile app behind a clean API.
Ingestion that cleans, chunks and indexes your content, and keeps the index fresh when documents change.
LLM-powered steps inside real processes — triaging emails, extracting fields from forms, routing tickets — with a person approving where it counts.
A test set of real questions, answer-quality checks, prompt-injection defences and limits on what the model is allowed to say.
Simple questions go to small, cheap models; hard ones to larger models. You pay for capability only when it is needed.
Signs you need this
Tech stack
Model-agnostic by design: we pick the model per task and keep you free to switch.
Google GeminiOpenAI GPTAnthropic ClaudeQwenOpen-weight models
EmbeddingsBM25 keyword searchHybrid rankingpgvectorMySQL full-text
PythonFastAPIPHPLaravelWordPress plugins
Evaluation setsPrompt-injection testsUsage and cost loggingRate-limit fallbacks
How we work
We collect the questions people actually ask and the content that should answer them, and agree what “good” looks like.
A working assistant on a slice of your content, tested against your questions — so you judge it on evidence, not a demo.
Guardrails, access rules, fallbacks between models, logging and cost limits before anyone outside the team uses it.
Rolled out on your site, app or internal tools, with a handover session and simple admin controls.
We review unanswered and low-rated questions, add content and tune retrieval on a regular rhythm.
Proof
Live
A RAG assistant on a Nigerian law firm’s website that indexes its pages, articles and guides, combines BM25 keyword search with Gemini embeddings and answers with Gemini.
80%
A cost-aware memory agent that routes each question to the right model tier — 80% cheaper in testing than always using the largest model.
Ways to work
Clearly defined deliverables at a fixed price — ideal for websites, audits and well-understood builds.
See packages →For bespoke products and platforms: a short discovery, then a written proposal with milestones and a fixed or capped price.
Request a quote →Ongoing support, maintenance and improvement with a guaranteed response time and a set number of hours each month.
Discuss a retainer →Flexible time-and-materials help for troubleshooting, code reviews, consulting and team augmentation.
Book time →Questions
Can’t see your question? Ask us directly — you’ll get a straight, practical answer, even if it’s “you don’t need us for this”.
It is an AI assistant that looks things up before it answers. When someone asks a question, it searches your own content for the most relevant passages and gives only those to the language model, which writes the answer from them. The result is grounded in your material rather than in whatever the model happens to remember.
Any language model can, which is why we design against it. The assistant is told to answer only from retrieved content, to say so when it cannot find an answer, and to point people to a human. We test it against a set of real questions before launch and keep checking afterwards.
We use provider plans and API settings that do not train on your data by default, keep your documents in storage you control, and only send the passages needed to answer a question. For sensitive material we can restrict which content is indexed and who can ask about it.
Whichever fits the task and budget. We often use Google Gemini, OpenAI and Anthropic Claude models, and open-weight models where data must stay on your servers. The code keeps the model swappable, so a price change or a better model does not mean a rebuild.
Running costs depend on the number of questions, the length of your content and the models chosen. We estimate it during the pilot using your real traffic patterns, and design for low cost — caching, small models for simple questions and tight context windows.
Yes. The assistant sits behind an API, so the same brain can power a website widget, an internal tool, a mobile app or a messaging channel. We usually start with one channel and add others once it is performing well.
We build with the Nigeria Data Protection Act 2023 in mind: collecting only what is needed, telling users how their data is used, protecting logs and giving you control over retention. We are engineers, not your compliance advisers, so we work alongside whoever handles data protection for you.
A sample of the content the assistant should know, a list of questions people really ask, and someone who can judge whether the answers are right. That is enough to start a pilot.
Tell us what you’re building, or what’s broken. We’ll come back with questions, a suggested approach and the simplest sensible next step — no obligation.