Search by job, company or skills

Principal AI/ML Engineer Lead

  • Posted 7 hours ago
  • Be among the first 10 applicants

Job Description

Description

WHAT YOU'LL DO

  • LLM Engineering Standards Across Agent Pods
    • Define and own LLM engineering standards across all Agent Stream pods — agent framework conventions, prompt lifecycle standards, eval harness design patterns, guardrail implementation, and confidence threshold calibration methodology that all Senior AI/ML Engineers follow.
    • Own the Langfuse observability framework at platform level — instrumentation standards, trace validation patterns, eval pipeline design, prompt regression testing, and model version regression detection. Pod-level Langfuse usage is consistent because this role defines how it is done.
    • Standardise RAG pipeline architecture across pods — embedding strategy, vector database selection and management, retrieval strategy, reranking, and structured output design for financial document reasoning. Shared RAG infrastructure is your design.
    • Review and challenge agent design proposals from pod AI/ML engineers — raise the bar on prompt design, memory architecture, evaluation rigour, and production reliability across the stream.
    • Contribute to AI/ML hiring — define the technical bar for AI/ML engineers across the stream, participate in interviews, and ensure hiring standards are consistent across pods.
  • Memory Architecture Ownership
    • Own the memory architecture strategy for the AI Platform — designing the personalised, complex memory layer that agents depend on for continuity, context, and adaptive behaviour across sessions and users.
    • Design and implement the tiered memory architecture for agent workflows — working memory (in-context), episodic memory (past interactions), semantic memory (extracted facts and preferences), and procedural memory (agent instruction updates). Select and implement the right memory framework for each tier: LangMem for LangGraph-native flows, Mem0 for managed personalisation, Zep/Graphiti for temporal and knowledge-graph reasoning, or Letta for explicit OS-style memory management.
    • Define memory hygiene standards — extraction policies, deduplication, contradiction resolution, and forgetting policies for agents that write aggressively to memory at scale.
    • Own context window management strategy — how long-running agent workflows handle context pressure, when to compress, when to retrieve from memory, and how to maintain coherence across multi-step financial close workflows.
  • Platform AI/ML Contribution
    • Contribute hands-on to Platform Team AI capabilities — working with the Platform Architect on how RAG, memory, and eval infrastructure is exposed as shared platform services that agent pods consume.
    • Stay current on the LLM and agent engineering landscape — evaluate new frameworks, protocols, and tooling (MCP, A2A, new memory systems, emerging eval approaches) and bring informed recommendations to the Director of Engineering on what to adopt and when.
    • Identify and address systemic AI/ML quality gaps across pods — inconsistent eval practices, weak guardrail implementations, or memory architectures that will not scale.
Who You Are

  • Extensive experience in AI/ML engineering, LLM engineering, data science, software engineering, or a related technical discipline, including experience delivering production-grade AI or ML solutions.
  • Strong hands-on experience with production-grade LLM agent development, including LangChain, LangGraph, or similar agent frameworks.
  • Experience defining platform-level engineering standards, architecture patterns, reusable frameworks, or technical practices across multiple teams or product areas.
  • Strong understanding of prompt lifecycle management, including prompt versioning, rollback, environment-specific configuration, evaluation harnesses, and regression testing.
  • Experience with LLM evaluation and observability practices, including confidence threshold calibration, guardrail design, model regression detection, trace validation, and tools such as Langfuse or similar platforms.
  • Experience designing RAG pipeline architecture, including embedding strategies, vector database selection, retrieval strategies, hybrid search, reranking, and structured output design.
  • Experience designing or implementing agent memory systems, context management strategies, or long-running agentic workflows.
  • Strong data science foundation, including model evaluation, statistical reasoning, experimental design, and understanding of model behavior in non-deterministic or edge-case scenarios.
  • Production engineering experience with Python and related API frameworks, such as FastAPI or similar tools.
  • Experience with PostgreSQL, pgvector or similar vector database technologies, Docker, Kubernetes, and Azure OpenAI or equivalent LLM providers.
  • Ability to evaluate emerging AI/ML technologies, make informed architecture recommendations, and guide technical decisions across teams.
  • Strong communication, collaboration, technical leadership, and problem-solving skills.
  • Experience with Ragas or equivalent RAG evaluation frameworks preferred.
  • Experience with model fine-tuning or RLHF preferred.
  • Experience with financial close, Record-to-Report, accounting, or enterprise finance software preferred.

At our core, Trintechers stand committed to fostering a culture rooted in our core values – Humble, Empowered, Reliable, and Open. Together, these values guide our actions, define our identity, and inspire us to continuously strive for excellence in everything we do.

All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin or disability.

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 151731027