Search by job, company or skills

Generative AI Developer

Generative AI Developer

Sonata Software
Fresher
Not Disclosed
  • Posted 17 days ago
  • Over 50 applicants have applied

Job Description

Job Summary:

Key Responsibilities

  • Build and extend PDF text & layout extraction using PyMuPDF for text-based PDFs, and Claude Vision for scanned, multi-column, or complex pages, producing bounding-box-anchored regions.
  • Generate multilingual/cross-lingual embeddings (via Bedrock) to shortlist candidate PDF regions matching each English instrument item.
  • Integrate a vision-capable Claude model on Amazon Bedrock as a verifier to confirm the correct source region for each item, with confidence scoring and a small-model fallback chain; ensure unmatched or low-confidence items are flagged rather than guessed.
  • Build deterministic placement logic (AWS Lambda) that clones the English skeleton, injects matched text, and validates structural integrity — string IDs, ordering, and formatting tags.
  • Design and build the agent orchestration layer using AWS Agent Core / MCP Server, connecting extraction, retrieval, verification, and placement as distinct tools with clear interfaces.
  • Implement the extract retrieve verify place validate workflow and guardrails using AWS Strands (Agentic SDK), ensuring structure-locked output, zero invented text, and full source traceability.
  • Support the linguist review & approval tool, showing each output string alongside its PDF source region, and implement a per-string audit trail written to Amazon S3 / DynamoDB.
  • Integrate with the existing eCOA ingest / screenshot-review pipeline: accept English instrument JSON and translated PDFs as inputs, and emit a draft JSON in the required target schema.
  • Run validation across sample instruments and priority locales (including at least one non-Latin-script locale), and track/report placement accuracy and flag-rate metrics (target 90% correct placement, zero invented translations).
  • Participate in daily scrum and weekly governance calls with Sonata and Medidata stakeholders; track progress in Jira.

Required Skills & Experience

  • Hands-on experience with AWS Bedrock — invoking foundation models (Claude), including vision-capable models for document/image understanding.
  • Experience building agentic AI systems using AWS Agent Core / MCP Server (Model Context Protocol) and AWS Strands, or comparable agent orchestration frameworks (e.g., LangGraph, LangChain agents).
  • Strong Python development skills, including experience with PDF parsing libraries (PyMuPDF, pdfplumber, or similar).
  • Experience with embeddings-based retrieval / semantic search, including cross-lingual or multilingual embeddings.
  • Working experience with AWS Lambda, Amazon S3, DynamoDB, and CloudWatch.
  • Solid understanding of prompt engineering and LLM verification/guardrail patterns — confidence thresholds, hallucination avoidance, fallback chains.
  • Comfortable working with structured JSON schemas and schema validation.
  • Experience delivering in Agile/Scrum teams using Jira.

Preferred / Nice to Have

  • Experience in healthcare, life sciences, or clinical trials technology (eCOA, ePRO, EDC systems).
  • Exposure to localization/translation-adjacent workflows or NLP/linguistics-adjacent projects.
  • Experience with OCR or vision-based document understanding at scale.
  • Prior experience integrating AI/automation solutions with third-party enterprise platforms via APIs.

Key Skills

PyMuPDF

AWS Bedrock

pdfplumber

AWS Agent Core MCP Server

AWS Strands

embeddings-based retrieval

JSON schemas

Claude

About Company