Search by job, company or skills

Senior AI Engineer

  • Posted 2 hours ago
  • Be among the first 10 applicants

Job Description

About Metaforms

Market research runs on 30-year-old survey platforms and armies of specialists hand-coding questionnaires in proprietary languages. Metaforms is the agent layer that does that work. Every survey is a program — full of skip logic, piping, quotas, and loops — and a single wrong number in a client report is unrecoverable. Our AI agents write production survey code, QA live deployments, process and clean large structured datasets, configure analysis, and generate client-ready reports, so agencies like Dynata, Savanta, and Borderless Access ship more projects with far less friction.

  • 1,000+ surveys processed monthly
  • Serving Fortune 500 companies across the globe
  • Rapid month-over-month growth

We're Series A funded and scaling fast, aggressively growing our AI engineering team to build the next generation of production-grade AI agent systems.

The Role

We're hiring a Senior AI Engineer to own the design, development, and continuous improvement of the AI agent systems that power modern research operations.

This is a high-ownership, high-impact role at the intersection of applied AI and systems engineering. You'll work on genuinely hard problems: agent reliability at scale, long-context handling, cascading error mitigation, and evaluation infrastructure — like codegen agents that write in proprietary DSLs, computer-use agents that QA live deployments, data agents that clean tabular exports and configure multi-step analysis, and evals for outputs where correct is genuinely ambiguous. And you'll do it on a team that ships fast and treats quality as non-negotiable.

What You'll Own

Agent Harness and Architecture

  • Own the agent harness our production agents run on — the loop where agents plan, use tools, check their work, and recover from failures
  • Lead research and implementation for long-context handling and cascading-error challenges in multi-step agent pipelines
  • Drive context engineering strategy and experimentation frameworks across the team

Evaluation and Production Monitoring

  • Define structured rubrics for evaluating AI outputs on nuanced, ambiguous research tasks
  • Build continuous monitoring, tracing, and failure-mode analysis for agents in production — including the loop that turns production failures into test cases
  • Create tooling that lets domain experts refine and evolve the skill files, eval sets, and knowledge bases our agents consume

Reliability for High-Stakes Outputs

  • Build eval suites — regression sets, golden datasets, LLM-as-judge pipelines — that catch regressions before deploy
  • Develop evaluation datasets for DSLs, structured data transforms, and computed outputs to systematically find and close model weaknesses
  • Design human-in-the-loop and review workflows for outputs where a single wrong number in a client report is unrecoverable

What We're Looking For

Must-Have

  • Built and operated agentic systems in production — multi-step pipelines, tool use, codegen, computer-use, or data and reporting agents — not just prototypes
  • 4+ years of engineering experience, with at least 1 year focused on LLM/agent systems in production
  • Deep hands-on experience with frontier model APIs (Anthropic, OpenAI, Gemini), evaluation frameworks, and AI system optimization
  • Strong Python skills; Go or TypeScript a plus
  • Solid grasp of context engineering and evaluation methodology
  • Strong instincts for debugging complex, non-deterministic system failures
  • High ownership: you drive problems to resolution independently and pull others in when it matters

Nice to Have

  • Experience with LLM observability and eval tooling (Braintrust, Langfuse, LangSmith, Weave, promptfoo, or in-house equivalents)
  • Background in semantic parsing, DSLs, or structured-output generation
  • Prior work on computer-use or browser agents
  • Experience with human-in-the-loop agent workflows where proposals are reviewed before apply, or agents over large structured datasets

Why Metaforms

  • Work at the frontier of production AI: systems handling 1,000+ research projects a month, with the reliability bar that implies
  • A small, senior team where your decisions carry real architectural weight
  • Zero-bureaucracy culture: high autonomy, fast feedback loops, direct access to leadership
  • Well-funded and financially stable, with a clear roadmap and the runway to execute on it

Benefits

  • Full family health insurance
  • $1,000 USD annual learning and development budget
  • Dedicated mentor and coaching support
  • Free snacks and dinner at the office

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 151731999

Similar Jobs

Bengaluru, India

Skills:

KafkaTypescriptRESTGcpKubernetesPythonAWSagentic AILLMsGoGRPCAI-assisted development

Bengaluru, India

Skills:

.Net Core.NETSqlRabbitmqMicroservicesDeep LearningTensorflowKafkaMachine LearningAWSPytorchPythonAzureGcpRestful ApisAi

Bengaluru, India

Skills:

UipathApi ManagementAws ServicesPower AutomateFunctionsRpa ToolsLambdaDockerApi GatewayAWS BedrockAKSAzure OpenAIAI automation architecturesAzure servicesMicrosoft Power Platformcontainer orchestration

Bengaluru, India

Skills:

GoogleModel Context Protocol MCPLang Graphvector databasesLang ChainAnthropicGen AIOpen AILLM applicationsretrieval-augmented generation pipelinesLlama Indexhybrid searchAI ML systemsgraph-native vector indexes

Bengaluru, India

Skills:

DockerRest ApisKubernetesPythonLangChainAnthropic ClaudeCrewAIGoogle Geminisimilarity search infrastructureembedding modelsvector databasesAutoGenLangGraphLLM APIsmicroservices architectureMistralLlamaOpenAI GPT-4semantic searchLlamaIndexRAG architectures