Research Engineer, Agents for Security Operations
Build and train agents that do SOC work that can be checked. You write the code, run the experiments and own the result: the simulators they learn in, the training that shapes them, and the evaluations that decide what an agent may do on its own.
The role
Oryen automates security operations with agents whose work can be checked. Agents gather evidence from the SIEM, EDR and identity tools a SOC already runs, and a fixed set of rules, not the model, decides the verdict.
Your job will be to make those agents better at SOC work and to prove it: the environments they learn in, the scenarios they are tested on, the signal that shapes them, and the evaluations that decide what they may do on their own. You write the simulator, the training loop and the eval harness yourself, and your code runs in the product.
Making the agents better
Today the agents follow prompts and playbooks. We want them to learn: to find the decisive fact in fewer queries, to say “I don’t know” when the logs do not support a claim, and never to learn a shortcut to a wrong answer. They run inside the customer’s network on consumer-grade hardware, so the models are small and the intelligence has to come from training and evidence.
How you get there is open. Reinforcement learning, LoRA and other parameter-efficient fine-tuning, supervised fine-tuning on good traces, better retrieval, better evaluation — what matters is what moves accuracy and autonomy, not which method is fashionable this year. You are the one who argues for a choice and then shows whether it worked.
A large part of the job is simulating attacks and non-attack scenarios, the false alarms that make up most of a SOC’s day, analysing how the agent handles them, and turning that into training signal.
- Environment. A deterministic, replayable simulator of a SOC: alerts, the logs behind them, the tools the agent calls, realistic gaps and noise.
- Scenarios. Families of attack, benign, degraded and adversarial variants generated from real cases and ATT&CK techniques.
- Training signal. Correct verdict, correct facts, query cost, calibrated uncertainty — in whatever shape the method needs, whether that is a reward, a preference or a labelled trace. Find the shortcut the model would rather take, and close it.
- Evaluation. Held-out scenarios, replay against past cases, wrong-close rate as the headline metric. Evaluation is what decides how much an agent is trusted to do unsupervised.
Investigation is where this starts. Threat hunting and detection engineering come next, on the same loop.
What we are looking for
- You build, and you have trained something rather than only prompted it.
- You have built an environment or an evaluation and found the holes in your own benchmark.
- You know what a few billion parameters can and cannot do.
- Sceptical by default, clear in written English.
- Nice to have: SOC exposure, synthetic data or simulation work, formal verification, Czech.
How to apply
Send a CV and a link to one thing you built that we can read. Three conversations: an introduction, a deep dive into something you built, and a pairing session designing a scenario family.
We have deliberately not written down a seniority band, a location or a contract shape. Those depend on who you turn out to be, and we would rather work them out with you than filter you out with them.