Pythrust

Personalisation and recommendations

A recommender is not a model. It is a funnel with a model near the end.

Personalisation that narrows your entire catalogue to the few things worth showing this user right now. Built on your own events, and measured against a holdout so you can prove it lifted anything.

Personalisation that adapts, evolves,
and responds to every user in real time.

User Journey Audit

We look at what your users actually do rather than what the personas say, and find the surfaces where a better ordering would change the outcome.


Experience Intelligence Design

Adaptive flows and recommendation logic driven by behavioural signals and real time context, including what to do when there is no signal yet.


Model Integration

Retrieval, ranking, and clustering models wired into your product, serving inside a latency budget your product can actually absorb.


Custom Recommendation Engines

When an off the shelf engine cannot express what matters in your domain, we build one that can, on your own definition of a good result.


Data Infrastructure

Clean event pipelines and user models, because personalisation is downstream of instrumentation and no model recovers from bad events.


Team Enablement

We leave your team able to read the metrics, run the tests, and change the strategy without calling us first.

What changes once it works

  • Show each user the part of your catalogue they would otherwise never reach.
  • Keep people returning by adapting to what they do rather than what they signed up as.
  • Lift conversion by ordering for intent, at the moment intent is highest.
  • Prove the difference with a holdout, so the lift is a measurement rather than a belief.
Talk to our experts

The pipeline

How a catalogue becomes the five things a user sees.

The ranking model gets the attention. Most of the work, and most of the failure, happens in the stages either side of it.

Exhibit 1

The recommendation funnel, from everything you could show to the handful you do.

  1. 01

    Catalogue and events

    Everything you could show, and everything users have done.

    Everything
  2. 02

    Candidate generation

    Cheap and fast, tuned for recall. Missing a good item here is unrecoverable.

    Thousands
  3. 03

    Filtering

    Availability, eligibility, region, and anything you have said never to show.

    Hundreds
  4. 04

    Ranking

    The model scores what is left for this user, in this context, right now.

    Dozens
  5. 05

    The surface

    What is rendered, plus a slice held back to keep learning.

    What fits on screen

What the user does with those results becomes tomorrow's signal, and feeds straight back into stage one.

Every recommendation you show becomes training data for the next one. Without deliberate exploration, a recommender quietly narrows to what it already believed.

Stages two to five are what we build. The catalogue and the raw events stay in your systems throughout.

Our approach

  1. Behavioural Data Understanding

    We check whether your events can actually support personalisation before anyone builds a model.

    You leave with

    • An audit of the signals you have
    • The ones you are missing
    • A first use case that is feasible
  2. Model-Powered Personalisation

    Candidate generation and ranking built on your data and your definition of a good result.

    You leave with

    • A working pipeline end to end
    • Offline evaluation you can rerun
    • A cold start path for new users
  3. Seamless Product Integration

    It serves inside your product within a latency budget, with business rules you control.

    You leave with

    • Live on one real surface
    • Rules and overrides in your hands
    • A holdout group in place
  4. Continuous Learning & Optimisation

    Measured against the holdout, tuned from real behaviour, with exploration kept deliberate.

    You leave with

    • Lift reported against a control
    • Exploration tuned, not accidental
    • A monthly review of what moved

How we engage

Three ways to work with us.

Which one fits depends on whether you are checking your data can support this, building the first surface, or proving the lift.

  • Signal audit

    Best when you want personalisation but are not certain your event data can carry it yet.

    What is included

    • An audit of your events, catalogue, and user model
    • What signals exist against what the use case needs
    • An instrumentation plan for the gaps
    • A feasible first surface, with the reasoning

    Fixed price for a defined window.

  • First surface build

    Best when the data holds up and you want one recommendation surface live and measurable.

    What is included

    • Stages two to five of Exhibit 1, built on your data
    • Integration into one real surface in your product
    • A holdout group and a measurement plan agreed upfront
    • Handover with the pipeline and evaluation in your accounts

    Fixed price, agreed after the audit. Billed against milestones.

  • Improve and expand

    Best once something is live and the job becomes proving lift and extending it.

    What is included

    • Online tests against the holdout, reported honestly
    • Cold start and exploration tuned rather than left to chance
    • New surfaces added as earlier ones prove out
    • A monthly review of what moved and what did not

    Monthly retainer. Rolling term, cancellable with notice.

How we arrive at a number

What this costs depends almost entirely on the state of your event data, which is why the audit comes first and is priced on its own. A 30 minute call at no charge, then a written scope with the work broken down and a fixed number against it. If the audit says your data cannot support this yet, we will tell you that rather than sell you the build.

Get a scope and a number

Fit

Whether this is right for you.

A good fit if

  • You have real event volume: views, plays, clicks, or purchases
  • There is more catalogue than any one user could reasonably browse
  • You can hold out a control group, so lift can actually be proven
  • You care about measured lift rather than about having the feature
  • Someone can define what a good recommendation means for your business

Probably not a fit if

  • You have a few dozen items, where good search and sorting would serve better
  • Event tracking is broken or nobody internally trusts the numbers
  • A control group is out of the question, which makes lift unprovable
  • You need it fully accurate on day one with no cold start period
  • The actual goal is being able to say the product uses AI

Common questions

Answered before you have to ask.

  • How much data do we need before this is worth doing?

    Less than most people assume, but it has to be clean and it has to be real interactions rather than page views. Thousands of genuine events across a few hundred items is usually enough to beat a popularity baseline. The audit exists to answer this for your data specifically, before you commit to a build.

  • What happens with new users and new items?

    Cold start is a design decision, not an edge case, so we plan for it explicitly. New users get context and popularity based results that improve within a session. New items get deliberately shown to a small slice of traffic so they can earn signal rather than being invisible forever because they have none. Skipping this is why catalogues ossify.

  • How will we know it actually worked?

    A holdout group that never sees personalised results, and a metric agreed before launch rather than chosen afterwards. Offline evaluation guides the build, but only an online test against a control tells you whether it lifted anything. If you cannot run a holdout, be sceptical of anyone reporting lift, including us.

  • Will it just keep recommending the popular items?

    It will if nobody stops it, because popular items get shown, get clicked, and become more popular. That feedback loop is the standard failure of a recommender. We counter it with a deliberate exploration budget and by measuring catalogue coverage alongside conversion, so narrowing shows up as a number rather than as a complaint from your merchandising team.

  • Can we control what it shows?

    Yes, and you should. Stage three of Exhibit 1 exists for exactly that: hard rules for availability, region, eligibility, and anything you never want surfaced, applied before the model gets a say. Business rules override the model rather than competing with it.

  • Do we need a data science team to run it?

    Not to operate it. The rules, thresholds, and exploration budget are configuration your product team can change, and the dashboards are built to be read without a statistics background. You need us, or someone like us, when the model itself needs to change, which is a far less frequent event than people expect.

0/1000