AI Evaluation Engineer
Brilliant Systems · Lahore, Lahore District; Punjab, Pakistan
FULL TIME
Job Description
Decide what done means. If we cannot score it, we have not specified it, and we certainly cannot sell it.
We sell fixed-price phases, so we need a defensible answer to whether a thing is finished. For deterministic software that is a test suite. For a system that is probabilistic by design it is an evaluation harness, and building good ones is a speciality. This role owns that speciality across client engagements and our own products.
What you will do
- Turn client requirements into measurable evaluation sets, including the adversarial cases nobody asked for.
- Build the harnesses, and the tooling around them, so any engineer can run and read them.
- Track drift in production and raise it before a client does.
- Work on BotUp, where every worker is reviewed and every run is metered and logged, so evaluation is part of the product rather than internal tooling.
- Push back early on scope that cannot be measured, while it is still cheap to change.
What we need from you
- You have built evaluation for a machine learning or LLM system that other people then relied on.
- Statistical literacy: you know what a small sample can and cannot tell you.
- Strong Python, and a taste for tooling other engineers actually adopt.
- Scepticism. This role only works if you are willing to report a result nobody wants.
Useful but not required
- A testing or QA engineering background before moving into machine learning.
- Experience with human-in-the-loop annotation at any scale.
- You have written about evaluation somewhere public.
Details
| Company | Brilliant Systems |
| Location | Lahore, Lahore District; Punjab, Pakistan |
| Type | FULL TIME |
| Niche | tech |
