Lead Software Engineer - Agent Safety

United States Digital Space LLC

Berlin

Vor Ort

EUR 120.000 - 170.000

Vollzeit

14 Tage+

Erhalte mehr Antworten von Arbeitgebern

Versende in nur wenigen Minuten einen passgenauen Lebenslauf.

Zusammenfassung

United States Digital Space LLC in Berlin is seeking a Lead Software Engineer to head the newly formed Agent Safety/Evals team. You will define technical strategy for evaluating models, enforce safety guardrails, and govern AI behavior across internal platforms and workflows.

You’ll partner with the AI Foundations team to ensure rapid AI adoption never compromises security, determinism, or quality, while building scalable evals, guardrails, and observability for internal use.

Qualifikationen

  • Proven technical leadership in complex, ambiguous domains.
  • Strong fundamentals in secure, end-to-end software design and testing.
  • Experience evaluating AI outputs and building reproducible evals.

Aufgaben

  • Define the technical strategy for evaluating models and enforcing safety guardrails.
  • Develop scalable evaluation harnesses tailored to internal workflows.
  • Implement governance for AI models operating within internal tools and platforms.
  • Drive AI security, observability, and incident response for internal agents and harnesses.
  • Lead verification and red-teaming to identify vulnerabilities in internal AI implementations.

Kenntnisse

Technical leadership
Python programming
AI evaluation & safety
Security & governance
AWS experience
Communication skills

Tools

AWS
Datadog
Django
LangChain
Pydantic AI
LiteLLM

Jobbeschreibung

Help us use technology to make a big green dent in the universe!

the company powers some of the most innovative global developments in energy.

We’re a technology company focused on creating a smart, sustainable energy system. From optimising renewable generation, creating a more intelligent grid and enabling utilities to provide excellent customer experiences, our operating system for energy is transforming the industry around the world in a way that benefits everyone.

It’s a really exciting time in energy. Help us make a real impact on shaping a better, more sustainable future.

the company is the operating system for the energy transition. We help energy companies, utilities, and system operators transform how they operate so the world can move faster towards a zero-carbon future. We build technology that solves real, messy problems at scale - across data, software, and increasingly AI. Our teams move fast, take ownership, and care deeply about impact.

AI is a key investment area for the company Technologies as we look to expand our existing capabilities. A crucial part of this is broadening the foundational infrastructure to enable teams across the organisation to use AI effectively to accelerate our mission.

You’ll work in the Agent Safety/Evals team, a new sub-team within AI Foundations. We build the shared platforms, harnesses, and guardrails that enable engineering and product teams to safely, reliably, and deterministically use machine learning and generative AI agents for internal systems and workflows across the business. This is a delivery-focused team that sits at the intersection of engineering and delivery, focusing on empowering our internal users.

The Role

We’re hiring a Lead Software Engineer to head our newly formed Agent Safety/Evals team. As AI agents take on more autonomous tasks across the company's internal workflows, this role is critical to ensuring they do so securely and predictably.

This is a leadership role focused on constraining our internally facing AI agents to an expected operation space and guaranteeing system reliability. You will define the technical strategy for how we evaluate models, enforce safety guardrails, and govern AI behavior across internal platforms, skills, and harnesses. You’ll work closely with the broader AI Foundations team and engineers across the company to ensure that our push for rapid AI adoption in internal tooling never compromises on security, determinism, or quality.

What you'll own
  • Lead the technical direction for internal agent safety: Design and implement systems focused on the reliability, security, and determinism of LLMs and autonomous agents, constraining them strictly to expected operational spaces within our internal ecosystems.
  • Build robust evaluation frameworks: Develop scalable harnesses and evals tailored to internal workflows and skills. You will be responsible for asking the right types of questions about the quality and reproducibility of our evals, while engineering the systems to measure them robustly.
  • Implement guardrails and governance: Create and enforce pre- and post-generation guardrails, managing the overarching governance of AI models operating within internal tools and platforms.
  • Drive AI security and observability: Build out dedicated auth/permissions for internal AI agents, establish deep monitoring/observability pipelines, and define incident response protocols for AI-specific anomalies.
  • Verification and Red Teaming: Lead continuous verification efforts and red teaming exercises to proactively identify vulnerabilities, prompt injections, or unpredictable behaviours in our internal AI implementations.
  • Operate in AWS: Deploy, run, and support high-throughput, low-latency safety services for internal use cases; make sensible architecture/cost tradeoffs; partner effectively with platform/techops/security stakeholders.
What you bring to the party
  • Strong technical leadership: Proven experience leading technical initiatives or teams, capable of setting the technical vision for complex, ambiguous domains.
  • Deep software engineering fundamentals: Senior/advanced capability in designing secure components end-to-end, testing thoroughly, and reasoning heavily about system design, concurrency, and architecture tradeoffs. (Python preferred).
  • Expertise in AI Evaluation and Safety: A highly critical thinker who understands the nuances of LLM behavior. You must be comfortable interrogating the quality of AI outputs and deeply experienced in building harnesses to measure it reproducibly.
  • Security and Governance mindset: Experience with threat modeling, authentication, red teaming, or building guardrails for internal systems, tooling, or platforms.
  • Cloud experience (AWS): Comfortable running internal services, owning reliability/scalability, and collaborating with platform/techops/security partners.
  • Excellent communication: Able to synthesize complex safety and eval metrics into actionable insights for both technical and non-technical stakeholders.
What Success Looks Like
  • Confidence in Internal Tooling: Engineers across the company can deploy new autonomous models, AI skills, and internal agents with total confidence, knowing they are bound by reliable safety guardrails and strict operational constraints within our internal ecosystems.
  • High-Quality Internal Evals: You have established a culture and infrastructure of robust, highly reproducible evaluation metrics that accurately reflect the real-world performance and safety of our internal-facing AI tooling and harnesses.
  • Resilient Internal Infrastructure: AI security, monitoring, and incident response are treated as first-class citizens, ensuring that any erratic agent behavior in our internal platforms is immediately caught and mitigated before it affects broader operations.
  • Technical Leadership: Strong collaboration across AI Foundations helps accelerate secure internal AI adoption; you are actively mentoring engineers and shaping the company's internal AI governance strategy.
  • Bonus points
  • Experience with specific LLM evaluation and safety frameworks (e.g., Inspect AI, Ragas, OpenAI Evals, NeMo Guardrails).
  • Django experience and strong backend engineering patterns (security, performance, maintainability).
  • Experience with Datadog for complex observability, tracing, and monitoring in AI environments.
  • Familiarity with foundational AI engineering tooling like Pydantic AI, LiteLLM, or LangChain.

the company is a certified Great Place to Work in France, Germany, Spain, Japan, Australia, and USA. In the UK we are one of the Best Workplaces on Glassdoor with a score of 4.5, and in Germany we rate 4.7 on Kununu as a Top Company. Check out our Welcome to the Jungle site (FR/EN) to learn more about our teams and culture.

*Are you ready for a career with us? We want to ensure you have all the tools and environment you need to unleash your potential. If you have any specific accommodations or a unique preference, please contact us at hr@unitedstatesdigital.space and we'll do what we can to customise your interview process for comfort and maximum magic!*

*Studies have shown that some groups of people, like women, are less likely to apply to a role unless they meet 100% of the job requirements. Whoever you are, if you like one of our jobs, we encourage you to apply as you might just be the candidate we hire. Across the company, we're looking for genuinely decent people who are honest and empathetic. Our people are our strongest asset and the unique skills and perspectives people bring to the team are the driving force of our success. As an equal opportunity employer, we do not discriminate on the basis of any protected attribute. We consider all applicants without regard to race, colour, religion, national origin, age, sex, gender identity or expression, sexual orientation, marital or veteran status, disability, or any other legally protected status.*

*Our (i)* *Applicant and Candidate Privacy Notice and Artificial Intelligence (AI) Notice**, (ii)* *Website Privacy Notice* *and (iii)* *Cookie Notice* *govern the collection and use of your personal data in connection with your application and use of our website. These policies explain how we handle your data and outline your rights under applicable laws, including, but not limited to, the General Data Protection Regulation (GDPR) and the California Consumer Privacy Act (CCPA). Depending on your location, you may have the right to access, correct, or delete your information, object to processing, or withdraw consent. By applying, you acknowledge that you’ve read, understood and consent to these terms*

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Hol dir deinen kostenlosen, vertraulichen Lebenslauf-Check.
oder ziehe deine Datei hierhin.
Similar jobs

Ähnliche Jobs, die dir auch gefallen könnten

Enterprise Account Executive, Digital Native Business - Munich
Enterprise Account Executive, Digital Native Business - Munich

United States Digital Space LLC • München

Hybrid
EUR 205.000 - 255.000
Staff Software Engineer - AI Platform Team
Staff Software Engineer - AI Platform Team

United States Digital Space LLC • Deutschland

Hybrid
EUR 70.000 - 90.000
AI Engineer (Germany)
AI Engineer (Germany)

United States Digital Space LLC • Berlin

Hybrid
EUR 70.000 - 120.000
Flexible PTO
Health insurance
Employee assistance programs
+4
Senior Forward Deployed Engineer
Senior Forward Deployed Engineer

United States Digital Space LLC • Berlin

Hybrid
EUR 90.000 - 130.000
Account Director, Digital Natives - DACH OpenAI Munich, Germany
Account Director, Digital Natives - DACH OpenAI Munich, Germany

Neura Market • München

Hybrid
EUR 120.000 - 190.000
Relocation assistance
Hybrid work model
Lead Forward Deployed Engineer
Lead Forward Deployed Engineer

United States Digital Space LLC • Berlin

Hybrid
EUR 90.000 - 160.000
Equity and cash compensation
Self-development budget
Hybrid work model
+1
AI Operations Lead
AI Operations Lead

United States Digital Space LLC • Berlin

Hybrid
EUR 95.000 - 150.000
Top-of-market equity & cash package
Self-development budget
Home office setup
+1
AI Deployment Engineer, Startups
AI Deployment Engineer, Startups

United States Digital Space LLC • Deutschland

Hybrid
EUR 75.000 - 90.000
Relocation assistance
Senior AI Software Engineer
Senior AI Software Engineer

United States Digital Space LLC • Berlin

Vor Ort
EUR 90.000 - 130.000
Senior AI Engineer - Agentic AI Evaluation Brain Team · Munich, Singapore ·
Senior AI Engineer - Agentic AI Evaluation Brain Team · Munich, Singapore ·

Resaro • München

Hybrid
EUR 90.000 - 130.000