On behalf of a boutique AI engineering consultancy and data lab: build the sandboxed environments, synthetic-data pipelines, and evaluation harnesses that frontier models are trained and tested against. A mid-level role for one of the fastest learners in the room — on-site in Yerevan, with full relocation sponsorship and covered housing.
ArgusRecruit is hiring on behalf of a boutique engineering consultancy and data lab operating at the frontier of artificial intelligence. The bottleneck in this field is no longer raw compute — it is high-signal, domain-specific, complex data. Our client's customers are the world's leading model providers, and its work is the infrastructure they rely on to go from 0 to 1 on their hardest post-training problems: complex execution environments, terminal-based benchmarks (in the vein of Terminal-Bench or OSWorld), programmatic synthetic-data pipelines, and bespoke evaluation harnesses. They don't scrape the open web — they build the “gyms” frontier models train in. This is a foundational hire on the core delivery engineering team.
The reality of this role — read this first
This is not a role where scope arrives pre-defined. One week the brief might be an evaluation harness for an agent resolving messy git merge conflicts in a headless Linux environment; the next, a synthetic-data generator that stress-tests multi-turn mathematical reasoning. Much of what you build has no precedent to copy — no StackOverflow thread, no textbook.
The company is looking for an engineer whose instinct at a technical wall is to open the paper, test three workarounds by the afternoon, read the raw stdout, and establish the ground truth themselves — not to flag it as blocked and wait for an architecture to be handed down. This role is likely the wrong fit if you measure your work in tickets closed rather than problems solved, treat understanding a system as a distraction from the job, or need a fixed six-month roadmap to feel settled.
What our client offers
The career fast-forward. Two years here compresses roughly six years of standard-industry learning, spent at the exact bottleneck of the AI field — enough to walk into a Tier-1 frontier lab afterward as a Data Systems Engineer.
Armenian work sponsorship. Full visa and legal relocation to Yerevan — one of the fastest-growing, safest, and most vibrant emerging tech hubs in the region.
A live-in engineering house. Housing is solved: a place in the company's live-in house in Yerevan, alongside a team of obsessed, high-agency peers.
Zero-cognitive-load compensation. With rent and utilities covered by the house, the base salary is calibrated for a comfortable, frictionless life in Yerevan — so your focus stays entirely on the work.
Real skin in the game. Generous early-stage equity in a boutique lab building high-value IP; your upside moves with the company, not a corporate banding chart.
Named exposure to the giants. Because the company works directly with model providers, your work is visible to the post-training engineering teams at the top AI labs — not buried in an internal repo.
Responsibilities
- Design and deploy programmatic evaluation environments — sandboxed OS environments, Dockerized test-beds, API mockers — that measure whether a model can actually perform complex tasks.
- Write the orchestration logic that generates, filters, mutates, and deterministically verifies high-quality synthetic datasets at scale.
- Build human-in-the-loop (HITL) tooling — internal workflows and custom annotation interfaces that let domain experts label complex data 10x faster.
- Translate ambiguous AI-lab requests (“a dataset that exposes where our model fails at spatial reasoning”) into concrete, executable software pipelines.
- Work directly against live client benchmarks, debugging the edge cases that have no documented solution.
Requirements
- A track record of shipping complex production systems. The specific language is not the point — what matters is a solid mental model of how computers work, since you'll increasingly direct AI agents to do the typing.
- Exceptional adaptability: dropped into a messy, half-documented open-source repository, you can grasp its core execution loop within an hour and bend it to a new purpose.
- High agency: you take a goal, map the intermediate steps yourself, and drive it to completion. You manage outcomes, not tasks.
- Systems pragmatism: async pipelines, concurrency, basic containerization (Docker), and talking to LLM endpoints without breaking rate limits or budgets.
- Zero dogma: you care whether the output is strictly correct, not whether the script that produced it used this month's trendy framework.
Nice to Have
- Hands-on work with LLM pipelines (LangChain, LlamaIndex, DSPy, or raw OpenAI/Anthropic orchestration loops).
- A conceptual grasp of the modern post-training stack (SFT, RLHF, DPO — and why synthetic-data collapse happens).
- Using, contributing to, or trying to beat open-source agentic benchmarks (SWE-bench, OSWorld, GAIA, etc.).
- A background in competitive programming, complex-systems debugging, or obsessive side-projects.