Personalisation and recommendations
Personalisation that narrows your entire catalogue to the few things worth showing this user right now. Built on your own events, and measured against a holdout so you can prove it lifted anything.
We look at what your users actually do rather than what the personas say, and find the surfaces where a better ordering would change the outcome.
Adaptive flows and recommendation logic driven by behavioural signals and real time context, including what to do when there is no signal yet.
Retrieval, ranking, and clustering models wired into your product, serving inside a latency budget your product can actually absorb.
When an off the shelf engine cannot express what matters in your domain, we build one that can, on your own definition of a good result.
Clean event pipelines and user models, because personalisation is downstream of instrumentation and no model recovers from bad events.
We leave your team able to read the metrics, run the tests, and change the strategy without calling us first.
The pipeline
The ranking model gets the attention. Most of the work, and most of the failure, happens in the stages either side of it.
Exhibit 1
Everything you could show, and everything users have done.
Cheap and fast, tuned for recall. Missing a good item here is unrecoverable.
Availability, eligibility, region, and anything you have said never to show.
The model scores what is left for this user, in this context, right now.
What is rendered, plus a slice held back to keep learning.
What the user does with those results becomes tomorrow's signal, and feeds straight back into stage one.
Every recommendation you show becomes training data for the next one. Without deliberate exploration, a recommender quietly narrows to what it already believed.
Stages two to five are what we build. The catalogue and the raw events stay in your systems throughout.
We check whether your events can actually support personalisation before anyone builds a model.
You leave with
Candidate generation and ranking built on your data and your definition of a good result.
You leave with
It serves inside your product within a latency budget, with business rules you control.
You leave with
Measured against the holdout, tuned from real behaviour, with exploration kept deliberate.
You leave with
How we engage
Which one fits depends on whether you are checking your data can support this, building the first surface, or proving the lift.
Best when you want personalisation but are not certain your event data can carry it yet.
What is included
Fixed price for a defined window.
Best when the data holds up and you want one recommendation surface live and measurable.
What is included
Fixed price, agreed after the audit. Billed against milestones.
Best once something is live and the job becomes proving lift and extending it.
What is included
Monthly retainer. Rolling term, cancellable with notice.
What this costs depends almost entirely on the state of your event data, which is why the audit comes first and is priced on its own. A 30 minute call at no charge, then a written scope with the work broken down and a fixed number against it. If the audit says your data cannot support this yet, we will tell you that rather than sell you the build.
Fit
A good fit if
Probably not a fit if
Common questions
Less than most people assume, but it has to be clean and it has to be real interactions rather than page views. Thousands of genuine events across a few hundred items is usually enough to beat a popularity baseline. The audit exists to answer this for your data specifically, before you commit to a build.
Cold start is a design decision, not an edge case, so we plan for it explicitly. New users get context and popularity based results that improve within a session. New items get deliberately shown to a small slice of traffic so they can earn signal rather than being invisible forever because they have none. Skipping this is why catalogues ossify.
A holdout group that never sees personalised results, and a metric agreed before launch rather than chosen afterwards. Offline evaluation guides the build, but only an online test against a control tells you whether it lifted anything. If you cannot run a holdout, be sceptical of anyone reporting lift, including us.
It will if nobody stops it, because popular items get shown, get clicked, and become more popular. That feedback loop is the standard failure of a recommender. We counter it with a deliberate exploration budget and by measuring catalogue coverage alongside conversion, so narrowing shows up as a number rather than as a complaint from your merchandising team.
Yes, and you should. Stage three of Exhibit 1 exists for exactly that: hard rules for availability, region, eligibility, and anything you never want surfaced, applied before the model gets a say. Business rules override the model rather than competing with it.
Not to operate it. The rules, thresholds, and exploration budget are configuration your product team can change, and the dashboards are built to be read without a statistics background. You need us, or someone like us, when the model itself needs to change, which is a far less frequent event than people expect.