Search Jobs

Search by job, company or skills

Applied Machine Learning Engineer (Model Layer)

Applied Machine Learning Engineer (Model Layer)

BruntWork
Fresher
  • Posted 4 hours ago
  • Be among the first 10 applicants

Job Description

Our client is building a simple, private AI chat platform designed to give users easy access to powerful open-source and leading AI models without requiring any technical setup. The product focuses on delivering high-quality AI responses while keeping the experience fast, affordable, and privacy-conscious.

The front-end and website are managed by a separate application team. This role is focused specifically on the AI model layer—the engine that powers the chat experience.

The team is small and early-stage, so this is an opportunity for a hands-on engineer to take significant ownership of a critical part of the product and work directly with founders and engineering stakeholders.

About the Role

As the Senior LLM / AI Model Routing Engineer, you will own the model layer from integration through production optimization. You will build the systems that determine which AI model should handle each request, evaluate model performance, and continuously improve the balance between quality, speed, reliability, and cost. This is a production engineering role, not a research-only position. The ideal candidate has experience taking LLM technology beyond prototypes and building reliable systems used by real users.

Key Responsibilities

LLM Integration & Model Management

  • Integrate open-source and third-party LLMs through APIs and inference providers.
  • Maintain the product's model lineup and evaluate new models as they become available.
  • Compare models based on quality, reliability, speed, cost, and suitability for different use cases.
  • Determine when models should be added, replaced, or retired.
  • Implement reliable model and provider fallbacks.

Model Routing

  • Design and maintain intelligent model routing logic.
  • Route requests to the most appropriate model based on factors such as task type, quality, latency, cost, and model behavior.
  • Continuously refine routing decisions using production data and user feedback.
  • Explore advanced approaches such as model cascades, ensembles, and multi-model systems when appropriate.

Evaluation & Optimization

  • Build practical systems for evaluating LLM output quality and routing decisions.
  • Establish and monitor key performance metrics, including quality, latency, reliability, and cost per request.
  • Analyze production data to identify opportunities for improvement.
  • Optimize prompts, model parameters, configurations, and fallback strategies.
  • Balance response quality with performance and operating costs.

Performance & Reliability

  • Monitor model performance, latency, usage, and inference costs.
  • Identify and resolve issues with models, providers, routing, and integrations.
  • Improve response speed and reliability as usage grows.
  • Implement appropriate logging, monitoring, error handling, and fallback mechanisms.
  • Maintain clear documentation for the model architecture and operational processes.

Collaboration

  • Work closely with the application engineering team to integrate the model layer into the product.
  • Collaborate on the connection between the AI, application, and policy layers.
  • Communicate technical concepts and trade-offs clearly to founders and non-ML stakeholders.
  • Provide data-driven recommendations on model strategy and product performance.

Requirements

  • Proven experience building LLM-powered applications in production.
  • Strong Python development skills.
  • Hands-on experience working with LLM APIs and inference platforms.
  • Experience with providers such as OpenAI, Anthropic, Together, Groq, or similar.
  • Strong understanding of LLM capabilities, prompt design, model evaluation, and production behavior.
  • Ability to make practical trade-offs between quality, latency, reliability, and cost.
  • Strong analytical and problem-solving abilities.
  • Comfortable working independently in a small, fast-moving startup environment.
  • Strong written and verbal English communication skills.

Preferred Qualifications

  • Experience with LLM model routing, multi-model architectures, ensembles, or AI gateways.
  • Familiarity with open-weight models such as Llama, Qwen, or Gemma.
  • Experience with LLM evaluation frameworks or automated evaluation pipelines.
  • Experience optimizing inference cost, latency, throughput, or reliability.
  • Experience with model observability and production monitoring.
  • Experience building privacy-focused or privacy-conscious AI products.
  • Previous experience working in an early-stage or lean engineering team.

The Ideal Candidate

We are looking for a builder who enjoys ownership and execution.

You don't need to be an academic AI researcher. You need to understand how modern LLMs work, know how to evaluate them in the real world, and be able to turn that knowledge into reliable production systems.

You are:

  • Pragmatic: You make smart trade-offs instead of over-engineering.
  • Data-driven: You measure performance and use evidence to make decisions.
  • Curious: You stay current with new models, providers, and AI techniques.
  • Proactive: You identify problems and opportunities without waiting for instructions.
  • Ownership-oriented: You take responsibility for the performance and reliability of your systems.
  • Collaborative: You can work effectively with engineers, founders, and non-technical stakeholders.

What Success Looks Like

  • Requests are consistently routed to the best model for the task.
  • Users receive high-quality, accurate, and relevant responses.
  • Latency remains fast and consistent as usage increases.
  • Cost per request is continuously optimized.
  • Model and provider failures are handled reliably.
  • New models are evaluated and integrated quickly when they add value.
  • Routing and model decisions are based on measurable production data.
  • The model layer is reliable, well-documented, and continuously improving.

Why Join

  • Own a critical AI layer: Take end-to-end ownership of the engine behind the product.
  • Work with leading LLMs: Continuously evaluate and work with new open-weight and provider models.
  • Make a direct impact: Your decisions will directly influence product quality, performance, and cost.
  • Work with a lean team: Collaborate closely with founders and have meaningful technical influence.
  • Remote & global: Work from anywhere with a flexible schedule.
  • Build for real users: Move beyond experimentation and build production AI systems.
  • Growth opportunity: Expand your technical ownership as the product and team scale.

Key Skills

Together

Model observability

Multi-model architectures

LLM integration

Model routing

Model cascades

Anthropic

Groq

Ensembles

LLM APIs

Inference platforms

Automated evaluation pipelines

Fallback mechanisms

Error handling

OpenAI

Prompt design

Model evaluation

About Company