Location: Remote, India
Schedule: Regular overlap with U.S. Eastern Time
We are looking for a Senior Applied AI Backend Engineer to help make our AI workflows reliable, secure, testable, and production-ready.
The Role
This is a backend-oriented applied AI engineering position.
You will build and maintain production LLM workflows, prompt systems, evaluation frameworks, backend services, database processes, and transcription integrations.
The primary focus is not frontend development or custom model training. The focus is turning probabilistic AI behavior into controlled, measurable, and defensible production systems.
Responsibilities
LLM Systems and Prompt Engineering
- Design and maintain multi-stage LLM workflows.
- Develop, test, version, and optimize system prompts, task prompts, evaluation rubrics, and structured outputs.
- Build reliable context-construction and prompt-orchestration processes.
- Implement model routing based on quality, latency, cost, availability, and task complexity.
- Design retry, fallback, timeout, repair, and graceful-degradation strategies.
- Understand the strengths and limitations of major commercial LLM providers.
- Determine when requirements should be enforced through prompts, deterministic code, validation, or database controls.
- Implement clear handling for insufficient evidence, conflicting data, malformed responses, and unsupported conclusions.
LLM Evaluations and Promptfoo
- Serve as a primary owner of Promptfoo implementation.
- Build and maintain automated LLM evaluation suites.
- Create representative datasets covering normal, difficult, ambiguous, and adversarial scenarios.
- Develop deterministic, semantic, schema-based, model-graded, and custom assertions.
- Compare prompts, models, configurations, and fallback strategies.
- Establish measurable quality thresholds and release gates.
- Integrate evaluations into CI/CD processes.
- Convert production failures into permanent regression tests.
- Identify whether failures originate in prompts, models, data, validation, evaluation design, or system architecture.
AI Security and Output Reliability
- Test and protect LLM workflows against prompt injection, indirect prompt injection, jailbreaks, system prompt disclosure, and data leakage.
- Treat resumes, transcripts, job descriptions, and other external content as untrusted input.
- Implement evidence grounding, source attribution, structured validation, and unsupported-claim detection.
- Design systems that preserve traceability between source information and AI-generated conclusions.
- Develop safeguards for hallucinations, inconsistent outputs, missing evidence, and model disagreement.
- Support human review or escalation where automated conclusions are not sufficiently reliable.
- Help ensure AI-generated information is appropriate for compliance-sensitive employment workflows.
Backend and Database Engineering
- Build and maintain secure backend services using TypeScript, Node.js, Deno, serverless functions, or comparable technologies.
- Develop reliable APIs, asynchronous processing workflows, and external service integrations.
- Work deeply with PostgreSQL or comparable relational database platforms.
- Design schemas, queries, constraints, functions, indexes, and transactional workflows.
- Implement secure multi-tenant data access and authorization controls.
- Build idempotent processes that protect against duplicate execution, race conditions, partial failures, and inconsistent state.
- Diagnose complex issues across backend code, databases, model providers, transcription services, and third-party APIs.
Database Migrations and Release Reliability
- Create and maintain version-controlled database migrations.
- Detect and remediate schema drift between environments.
- Plan backward-compatible changes, data backfills, and production deployments.
- Test migrations against representative data.
- Develop rollback or forward-remediation strategies.
- Validate database functions, permissions, backend services, secrets, and integrations after deployment.
- Investigate production issues that do not occur in development or testing environments.
Transcription Systems
- Support production speech-to-text workflows using platforms such as Deepgram, Whisper, AssemblyAI, Azure AI Speech, Google Cloud Speech-to-Text, AWS Transcribe, or comparable services.
- Understand the differences between real-time streaming transcription and asynchronous batch transcription.
- Handle interim and final results, diarization, speaker mapping, segmentation, reconnects, duplicates, and partial failures.
- Design reliable batch-processing workflows with retries, idempotency, failure recovery, and completion handling.
- Preserve transcript source, version, speaker attribution, and processing status.
- Understand how transcription accuracy affects downstream LLM analysis, evidence extraction, and reporting.
- Help determine when a streaming transcript or a higher-accuracy batch transcript should be used.
AI Observability and Compliance
- Implement logging and tracing across LLM and transcription workflows.
- Track prompt versions, model versions, latency, token usage, evaluation results, retries, fallback events, and validation failures.
- Evaluate LLM observability tools such as Langfuse, LangSmith, Braintrust, Arize Phoenix, or comparable platforms.
- Support technical controls related to secure development, access management, change management, auditability, incident response, and SOC 2 readiness.
- Help translate compliance and responsible-AI requirements into testable engineering controls.
Required Qualifications
- Six or more years of professional software engineering experience.
- Strong recent backend engineering experience.
- Demonstrated experience building and operating production LLM applications.
- Advanced proficiency with TypeScript, JavaScript, Python, or another relevant backend language.
- Strong SQL and PostgreSQL experience.
- Experience building APIs, serverless functions, background workflows, or distributed backend services.
- Experience integrating commercial LLM providers.
- Strong understanding of prompt engineering, context construction, structured outputs, prompt chaining, and model limitations.
- Hands-on experience with Promptfoo or another sophisticated LLM evaluation framework.
- Experience creating evaluation datasets, assertions, regression tests, and release thresholds.
- Understanding of hallucinations, prompt injection, indirect prompt injection, data leakage, and AI application security.
- Experience implementing retries, model fallbacks, timeouts, validation, and graceful degradation.
- Experience with database migrations and production deployment processes.
- Strong understanding of authentication, authorization, secrets management, and multi-tenant SaaS security.
- Strong written and spoken English.
Strongly Preferred
- Advanced Promptfoo experience.
- Experience creating custom evaluation assertions, graders, or providers.
- Experience integrating LLM evaluations into CI/CD.
- Experience with managed PostgreSQL platforms and serverless backend functions.
- Experience with Deepgram, Whisper, or another production transcription platform.
- Experience with both streaming and batch transcription.
- Experience with LLM observability platforms.
- Experience in HR technology, healthcare, financial services, insurance, legal technology, or another compliance-sensitive industry.
- Experience supporting SOC 2 readiness or enterprise security reviews.
- Familiarity with SAST, DAST, application-security testing, and AI red teaming.
- Ability to understand and troubleshoot a React and TypeScript frontend when required.
What We Are Not Looking For
- A primarily frontend-focused engineer.
- A prompt writer without strong backend engineering skills.
- A general DevOps engineer without production LLM experience.
- A research-focused data scientist or machine learning engineer.
- Someone whose AI experience is limited to chatbot prototypes or simple API integrations.
- Someone who relies on manual prompt testing without measurable evaluation and regression controls.
- Someone who accepts AI-generated code or outputs without independent validation.
What Success Looks Like
Within the first six months, this engineer will:
- Become a technical owner of LLM evaluation and backend reliability practices.
- Expand automated Promptfoo evaluations and red-team coverage.
- Establish measurable quality and security release gates.
- Improve prompt and model versioning, routing, fallback, and validation.
- Strengthen database migration and deployment reliability.
- Improve observability across LLM, backend, and transcription workflows.
- Reduce unsupported conclusions, malformed outputs, and production-only failures.
- Help make AI outputs reliable, traceable, and defensible.
The ideal candidate understands that an LLM call is not a complete AI feature. Reliable AI requires strong prompts, backend engineering, evaluations, security controls, observability, and disciplined production operations.