Artificial intelligence
AI engineers, and agents that join your team.
Vetted specialist engineers placed onto your team across the Gulf, India and Singapore — the people who have already taken AI systems into production. And custom agents, trained on your own domain, that hold a seat beside them.
- Gulf · India · Singapore
- Engineers who have shipped it before
- Agents trained on your own record
The position
Nobody is short of ideas. They are short of people who have shipped one.
Almost every large enterprise now has a pilot. Far fewer have something a customer touches, an auditor accepts and a finance director will fund again next year. The distance between those two states is rarely closed by a better model. It is closed by the unglamorous work of turning a demonstration into a system that behaves the same way on a bad day as it did in the room.
That work has a shape. Retrieval that stays correct as the source documents change. Evaluation that catches a regression before a customer does. Cost and latency held inside a budget. Access control that survives an audit. A rollback path for the week a model provider ships an update you did not ask for. None of it is exotic. All of it is learned the first time by getting it wrong, which is why the second system is so much easier than the first.
The region makes the point sharper. Gulf states have put artificial intelligence at the centre of national programmes and are building sovereign cloud capacity to keep the data inside their borders, which turns AI from an IT initiative into a policy commitment with dates attached. India supplies the platform, data and machine-learning engineering depth the rest of the region draws on. Singapore is where regional headquarters set the governance standard the whole group is then held to.
So the answer has two sides, and most firms sell one. Engineers who have taken these systems into production before, placed onto your team and working to your standards. And agents built on your own documents, systems and precedent, scoped to a job you would otherwise be hiring for. Each side makes the other faster: the engineers build the agents, and the agents return the engineers to the work only people can do.
Specialist engineers placed onto the client team, and agents scoped to a seat beside them — one programme, staffed and built together.
Half one · the engineers
The roles that decide whether it reaches production.
Ten seats we are asked for most. Each one is filled by someone who has held it before on a system that went live — engaged as an individual, a squad, or a team we manage against your outcomes.
- 01
Machine learning engineer
Takes a model from a notebook to something that serves traffic, retrains on a schedule and degrades predictably rather than silently.
- PyTorch
- scikit-learn, XGBoost
- Feature stores (Feast, Tecton)
- Ray
- MLflow
- Training and inference cost control
- 02
LLM and applied AI engineer
Designs the system around the model — retrieval, tools, structured output, guardrails — and is the person who knows why it answered that way.
- LangGraph, LangChain
- Retrieval and chunking design
- Function calling and structured outputs
- Fine-tuning (LoRA, DPO)
- Prompt and context engineering
- Token cost and latency budgeting
- 03
Data engineer
Builds the pipelines everything above depends on, including the ones that have to be right on the day the source system changes shape.
- Spark
- dbt
- Airflow, Dagster
- Kafka
- Snowflake, Databricks
- Iceberg, Delta Lake
- 04
MLOps and ML platform engineer
Makes deploying a model as boring as deploying a service — versioned, tested, reversible, and not dependent on the person who trained it.
- Kubernetes
- Terraform
- CI/CD for models and prompts
- Model and feature registries
- Canary releases and rollback
- Drift and cost monitoring
- 05
AI infrastructure engineer
Owns the accelerator layer — scheduling, throughput, memory and spend — which is where most inference budgets are quietly lost.
- CUDA, NCCL
- vLLM, TensorRT-LLM
- Slurm and Kubernetes GPU scheduling
- Quantisation and batching strategy
- Distributed training topology
- Capacity and cost planning
- 06
Data scientist
Works out whether the thing is actually true before the business builds on it, and says so plainly when the data will not support the claim.
- Experiment and holdout design
- Causal inference
- Python, SQL
- Forecasting and time series
- Statistical validation
- Uplift and segmentation analysis
- 07
AI product engineer
Builds the surface a person actually uses — streaming, citations, review states, the correction path — so the system is trusted rather than merely correct.
- TypeScript, React, Next.js
- Streaming and partial-response UX
- Human-in-the-loop review interfaces
- API and tool design
- Product telemetry
- Evaluation-driven iteration
- 08
Evaluation and quality engineer
Builds the harness that decides whether a change is an improvement, which is the role that turns opinion about a model into evidence.
- Golden and adversarial datasets
- Ragas, DeepEval
- LangSmith tracing
- LLM-as-judge calibration
- Regression suites in CI
- Annotation and inter-rater agreement
- 09
AI security engineer
Attacks the system the way an outsider would — prompt injection, tool abuse, data exfiltration through retrieval — and closes what they find.
- Prompt injection and jailbreak testing
- OWASP Top 10 for LLM applications
- Tenant isolation and access control on retrieval
- Secrets and key management
- Red teaming and abuse cases
- Audit logging and data-residency controls
- 10
Applied research engineer
Brought in when the standard approach has run out — ranking that will not improve, a domain no general model handles, an evaluation nobody has defined yet.
- Paper reproduction and ablation studies
- Retrieval and ranking research
- Fine-tuning and preference optimisation
- Metric and evaluation-suite design
- PyTorch at a low level
- Writing findings a business can act on
How we vet
Five stages, and none of them is a keyword search.
Each stage can end the process. Nothing here is exotic — the discipline is in refusing to skip a stage because the CV is impressive or the deadline is close.
- 01
Evidence, not keywords
We start from what the engineer has actually run: which system, under what load, who else was on it, what broke, and what they changed afterwards. A CV listing a framework tells you someone was in the room. This stage is designed to find out whether they were the person holding the decision.
- 02
A scoped practical
A short exercise shaped like the real work — a retrieval design that has to stay correct as the documents change, a pipeline that has to fail safely, an evaluation harness for a task with no single right answer. Scoped to a sitting rather than a weekend, and never work we could use ourselves.
- 03
Reviewed by someone who has shipped it
The practical and the technical conversation are assessed by an engineer who has built the same class of system in production, not by a recruiter with a scoring sheet. For the narrower roles — accelerator infrastructure, model security, evaluation design — the reviewer is a specialist in that field, brought in for the review.
- 04
References, right to work, and the file
References from people who managed them or worked beside them on named work, not character references. Right to work, sponsorship route and notice period confirmed before the name reaches you, because a perfect candidate who cannot mobilise inside your window is not a candidate.
- 05
A trial period, if you want one
For most engagements you can take the engineer on a paid trial before committing to the full term, on terms written down at the start rather than negotiated later. We would rather carry that risk than argue afterwards about whether the fit was as described.
Five stages, each able to end the process, with the practical assessed by someone who has shipped the same class of system.
Custom AI agents
An agent is a colleague with a scope, not a chatbot.
A chatbot answers whoever turns up. An agent has a job description: a defined seat on a named team, a scope of work it may act inside, a list of systems it may touch, and a rule that says when it must stop and hand to a person. Written that way it stops being a technology decision and becomes an operating one — the same conversation you would have about a new hire.
What makes it useful is what it is trained on, and that part is yours: your policies, your decided cases, your templates, your reference data, the precedent your team already argues from. A general model knows the industry. An agent grounded in your own record knows how your organisation has actually decided things, and can show you the document it took the answer from.
Then it works a shift. It takes the volume that never reaches the top of the queue, prepares the cases a person will decide, and escalates the moment it leaves its scope, with its reasoning attached so the handover is a review rather than a restart. Every correction your team makes goes back into the evaluation set, which is how the second month is better than the first.
Intake, unattended handling, and the escalation rule that returns a case to a person with its reasoning attached.
The team we deploy
Five seats, and what each one is for.
Every agent is built the same way: a scope it may act inside, the systems it may touch, the evidence it must attach, and the rule that ends its authority. What changes between them is the job.
Breaks are found daily and explained slowly. Tally does the explaining, so the team spends its day on the items that are genuinely unexplained.
Runs unattended
- Matches statement lines to ledger entries across accounts, currencies and value dates, including partials and rebookings.
- Writes the break narrative in the format the controller expects, with both source records attached.
- Proposes the correcting entry and the account it belongs to, without posting anything.
- Carries unresolved items forward with their age and history, so nothing quietly restarts at zero.
Trained on your
- Your reconciliation policy and break-classification standard
- Historical breaks and how each was finally resolved
- Chart of accounts, entity structure and counterparty reference data
- Statement and ledger formats from your own systems
Banking and financial services
Requests arrive as scans, portal submissions and faxes, and a clinician waits while somebody works out what is missing. Intake works out what is missing first.
Runs unattended
- Reads each request against the applicable policy and lists the clinical evidence that is present, absent or contradictory.
- Orders the queue by clinical urgency and by how close each request is to being decidable, rather than by arrival time.
- Drafts the information request back to the provider, quoting the criterion that is unmet.
- Assembles the reviewer's packet — the note, the criteria and the relevant history — in one view.
Trained on your
- Your medical policy and clinical criteria library
- Prior determinations and the rationale recorded against them
- Your formulary, code sets and provider network data
- The correspondence templates your team already uses
Healthcare and life sciences
Rotating equipment signals distress long before it stops, but the evidence sits across several systems and nobody reads all of them. Sentry reads all of them, continuously.
Runs unattended
- Correlates vibration, temperature and process data against the maintenance history for the same tag.
- Separates a real deviation from a sensor fault or a planned process change before anyone is paged.
- Opens a single notification carrying the trend, the comparable past events and a recommended inspection window.
- Keeps a running integrity picture per asset, so turnaround scope is argued from evidence rather than memory.
Trained on your
- Your historian tags, alarm configuration and process limits
- Maintenance and inspection records for the same equipment
- Closed failure investigations and their findings
- Your permit, isolation and integrity standards
Energy, oil and gas
Citizens and businesses wait on replies that are, in substance, replies already written before. Precedent writes the first version, grounded in the policy and the prior decisions that govern it.
Runs unattended
- Classifies the enquiry, finds the governing clause and retrieves the closest prior decisions.
- Drafts the reply in Arabic and English, to the department's own template and terminology register.
- Marks every assertion with the clause or precedent it rests on, so review is a check rather than a rewrite.
- Flags where prior decisions conflict with each other instead of quietly picking one.
Trained on your
- Published policy, regulations and circulars currently in force
- The department's own decided cases and their reasoning
- Approved bilingual templates and terminology register
- The retention and disclosure rules that apply to the reply
Public sector
Buyers decide range with whatever data they had time to gather. Range gathers continuously, so the buying meeting starts from a fuller picture of what is selling, what is missing and who could supply it.
Runs unattended
- Tracks the gap between demand signals and what is actually listed, by store cluster and by channel.
- Reads supplier catalogues, specifications and compliance documents to shortlist candidates against a gap.
- Separates why a line underperformed — price, placement, availability or the product itself — instead of blending them.
- Watches for lines heading out of stock ahead of a promotion that would make it worse.
Trained on your
- Your sales, stock and space data by store cluster
- Supplier master data, specifications and compliance files
- Past range reviews and the reasoning behind each decision
- Your promotional calendar and pricing rules
Retail and consumer
Working together
The agent takes the volume. The person keeps the judgement.
None of these pairings is designed to remove someone from the team. Each row splits a working day along the line where a decision starts to need accountability.
Your people
Defines what a complete case looks like, and what may never be assumed.
The agent
Assembles the case: pulls the records, quotes the source line behind every claim, names what is missing.
Together
The specialist opens a file that is already complete, and spends the hour on the judgement instead of the gathering.
Your people
Handles the exceptions, the disputes, and anything with a person's outcome attached to it.
The agent
Clears the routine majority unattended, and routes everything outside its scope with the reason stated.
Together
Queue depth stops being the thing that decides who gets attention today.
Your people
Owns the final wording, and signs it.
The agent
Produces the first draft in the house format, grounded in precedent and cited back to it.
Together
Review replaces authoring, and the house voice stops depending on who happened to be free.
Your people
Sets the thresholds, and decides what is worth waking someone for.
The agent
Watches continuously, correlates the signals, and opens a single issue rather than a wall of alerts.
Together
Attention goes to the problem instead of to sorting the noise that announced it.
Your people
Corrects the agent when it is wrong, and says why it was wrong.
The agent
Records every correction as a labelled case in the evaluation set, with the original reasoning kept.
Together
A correction becomes a regression test, so the same mistake is caught before it can be made twice.
A working day split along the line where a decision starts to need accountability: the agent prepares, the person decides.
Agents by industry
Every industry has work that never reaches the top of the queue.
Not a menu of products. These are the seats we are asked to fill most often, by sector — each one scoped, trained on the client’s own record, and answerable to a named person on the team.
Banking and financial services
Most of the cost sits in review queues — files that must be assembled, checked and explained before anyone is allowed to decide anything.
- Refreshhands off Any change in beneficial ownership or a screening hit that survives triage goes to a compliance officer.
Assembles the periodic review pack for an existing corporate relationship — ownership, screening and adverse-media checks — with every finding cited to its source.
- Narrativehands off Every escalation judgement and every regulatory filing stays with the investigator.
Writes the first-pass narrative for a transaction-monitoring alert, with the account history and the comparable prior alerts attached to it.
- Memohands off Ratings, conditions and the recommendation itself are set by the credit analyst.
Drafts the credit committee memo from spreads, filings, covenant history and the relationship team's own notes.
Healthcare and life sciences
Clinical and scientific time is consumed by documentation that is necessary, repetitive and unforgiving of error.
- Encodehands off A certified coder confirms or changes every code before anything reaches billing.
Proposes diagnosis and procedure codes from the clinical note, quoting the line of the record each code rests on.
- Cohorthands off A study coordinator confirms eligibility and makes every approach to a patient.
Reads protocol eligibility against de-identified records and produces a screening shortlist showing which criteria each candidate meets and misses.
- Vigilhands off Assessment, coding sign-off and reporting decisions remain with the safety physician.
Turns adverse-event reports arriving as free text across several channels into consistent structured cases, with seriousness and causality prompts raised.
Energy, oil and gas
Data arrives from the field faster than the people who understand it can read it, and the documentation burden is a safety obligation rather than administration.
- Permithands off Approval is always a signature from the permit authority, never an automated pass.
Checks permit-to-work applications against isolation rules, simultaneous-operations conflicts and certification currency before they reach the authority.
- Handoverhands off The outgoing supervisor edits, signs and owns the handover as issued.
Builds the shift handover from operator logs, alarm history and open maintenance notifications, ordered by what the incoming shift has to act on.
- Scopehands off Scope freeze and every deferral decision sit with the turnaround manager.
Reconciles inspection findings, deferred work and anomaly reports into a turnaround scope ranked by outage impact and access constraint.
Telecom and media
The network and the commercial catalogue both generate more exceptions than the teams reading them can absorb in a shift.
- Correlatehands off Field dispatch and customer notification decisions stay with the network operations lead.
Groups alarms across network elements into a single probable-cause ticket, with the affected services and customers already listed.
- Tariffhands off Anything needing a goodwill adjustment or a contract variation goes to a supervisor.
Answers frontline questions on plan rules, migrations and eligibility with the governing clause quoted rather than paraphrased.
- Clearancehands off Any ambiguous, disputed or lapsed right goes to legal before the asset can be scheduled.
Checks a media asset's rights window, territory and platform before it is scheduled, and flags what is about to expire.
Manufacturing
The knowledge that keeps a line running is real but scattered — across drawings, maintenance history and the people who have been there longest.
- Roothands off Root-cause conclusions and corrective actions are set and signed by the quality engineer.
Drafts the non-conformance investigation from inspection records, process data and every prior occurrence of the same defect.
- Cataloguehands off Anything safety-critical, or anything it cannot match confidently, goes to the maintenance planner.
Matches an unlabelled spare-part request to the catalogue using drawings, the equipment hierarchy and past purchase history.
- Changeoverhands off The planner commits the schedule; the agent never releases an order.
Sequences the production plan against material availability, tooling and changeover cost, and shows what each option gives up.
Retail and consumer
Decisions are made weekly on data that took most of the week to assemble, so the analysis arrives just as it stops being useful.
- Lifthands off Investment decisions for the next cycle sit with the category buyer.
Explains what a promotion actually moved by store cluster, separating price effect from availability and placement.
- Listinghands off Any failed compliance check goes to the technologist, and no listing goes live unapproved.
Checks a new supplier's product data, specifications and compliance documents against listing rules before anything enters the range.
- Returnhands off Product withdrawal and any supplier claim is raised by a person, on the record.
Reads free-text return reasons and flags sizing, quality and description defects to the buying team with worked examples.
Public sector
Caseloads grow faster than headcount, and consistency between officers matters as much as speed does.
- Scorehands off The award recommendation is made by the evaluation panel, on the record.
Scores tender submissions against the published criteria and shows the evidence sitting behind every score.
- Replyhands off Anything novel, appealable or marked sensitive is written by the case officer instead.
Drafts replies to citizen and business enquiries in Arabic and English, grounded in the policy in force and cited to it.
- Renewalhands off Refusals and any attached conditions are decided only by an authorised officer.
Checks licence and permit renewals for completeness and consistency with prior decisions before an officer opens the file.
Logistics and supply chain
Margin leaks through documentation errors and through exceptions found too late for anyone to do anything about them.
- Declarehands off Anything the broker must certify, and any classification dispute, goes to a person first.
Assembles customs declaration packs and checks classification, origin and valuation against the documents actually behind them.
- Diverthands off The recovery decision, and anything that changes a customer commitment, is made by the control tower lead.
Watches shipments against plan, opens the exception before the customer notices it, and puts a costed recovery option beside it.
- Demurragehands off Disputes are filed, and settlements agreed, by the account manager.
Reconciles carrier invoices against contracted free time and terminal timestamps, and prepares the dispute pack with the evidence attached.
The stack
Built on what teams actually run.
We are not tied to a platform, and we will not pretend a preference is a principle. Our engineers work in the stack you have chosen — and will tell you plainly when a piece of it is the reason the thing is slow.
Deployment
Everything runs inside your cloud account, your VPC and your identity provider. Your data does not leave your boundary to make an agent work.
Agent frameworks and orchestration
How work is decomposed, sequenced and made to survive a failed step. Most of the engineering in an agent lives here, not in the prompt.
- LangGraph
- LangChain
- LlamaIndex
- CrewAI
- AutoGen
- Semantic Kernel
- Temporal
- Model Context Protocol
- Azure AI Foundry
Models and serving
The models themselves, and the layer that serves them at a latency and a cost you can defend. Model choice is a per-task decision, not a corporate one.
- Claude
- GPT
- Gemini
- Llama
- Mistral
- vLLM
- NVIDIA Triton
- Amazon Bedrock
- Vertex AI
Retrieval and memory
Grounding an answer in your own record, and remembering what has already been established. The difference between an agent that quotes your policy and one that improvises around it.
- pgvector
- Qdrant
- Weaviate
- Pinecone
- Milvus
- Elasticsearch
- Redis
- Neo4j
Data and pipelines
The unglamorous layer everything above depends on. If the pipeline is late or wrong, no model choice repairs it.
- Databricks
- Snowflake
- dbt
- Apache Airflow
- Apache Kafka
- Apache Spark
- Ray
- Apache Iceberg
- Kubernetes
Evaluation, safety and observability
Whether it works, whether it is safe, and how you would know if it had quietly stopped working. This group is what decides whether a pilot is allowed to leave the pilot.
- MLflow
- Weights & Biases
- LangSmith
- Ragas
- DeepEval
- OpenTelemetry
- Arize Phoenix
- NeMo Guardrails
- Great Expectations
Systems of record, retrieval and memory, orchestration, model serving — and the evaluation and observability layer that sits across all of it.
Onsite where the programme requires presence, engineered from India where the supply is deepest, governed to the standard regional headquarters are held to.
FAQ
Questions we are asked first.
Tell us where it is stuck.
A first conversation is a working one, not a pitch. What you are building, what has stalled, and whether the answer is people, an agent, or both.