Tillbaka till alla jobb

AI Engineer (RL Environments)

Hashlist · Helsinki

Publicerad 07.09.2026 klo 03.00

Jobbeskrivning

Would you like to operate at the frontier of AI evaluation, post-training, and model improvements? We are now expanding our core Hashlist AI research team to support the creation of domain-specific RL environments for constraint-based embedded programming & complex enterprise engineering workflows. What you will do: Build and automate our platform for creating RL environments Construct simulated worlds and explore data shapes that expose meaningful model failure modes across the embedded coding domain & related enterprise workflows Turn AI training objectives into concrete data and evaluation specifications Build the reward layer + reward-hacking mitigation Run rollouts at scale. Hundreds of sandboxed attempts per task in parallel Fine-tune open-source models: Before-and-after fine-tunes, failure reports, and dataset exports Communication with our clients (OEMs & AI Labs) on specific research or fine-tuning questions they have Skills needed: Production machine learning depth. Python and PyTorch, reinforcement learning training with TRL, verl, SkyRL or OpenRLHF, PPO, GRPO or DPO, LoRA fine-tuning, and inference with vLLM or SGLang. Training or evaluation systems, including at least one RL environment you built end-to-end and trained a model against. Comfortable with the infrastructure around it, e.g Docker, or similar containerization tools, to design and monitor systems at scale Experience with harness & agentic optimisation for evals Strong familiarity with common reinforcement learning algorithms and methods, especially with respect to post-training LLMs High level of personal drive, motivation, and good communication skills. Bonus: You have published an environment, benchmark or evaluation harness we can look at. You have worked with embedded, safety-critical or other physical engineering software. Company benefits Competitive compensation + meaningful equity Central office in Helsinki Lunch benefit Be a part of a quickly scaling tech company working directly with model providers
Logga in gratis for att spara det har jobbet.