Who we are:
LEXI is a Silicon Valley start-up in the Enterprise IoT space that has already raised several million dollars from a long list of experienced angel investors.
In addition, the company's executive team includes many of the top IoT executives in the industry, such as multiple board members of the key wireless protocol standard setting bodies, product managers with over 70 IoT patents to their name and sales executives who have took their former company from $0 to $200M/annual revenue in 5 years (in IoT no less).
We are consistently winning major customers against our competitors, whose market caps start in the hundreds of millions of dollars and go into the billions. In short, you will be joining a winning team that is poised to grow the company into the unicorn ranks in just a few short years.
As with all our team members, you will be granted shares to participate in this huge growth trajectory.
Who we want:
We are looking for stellar Senior DevOps Engineer to join our team in Mumbai. We are looking for a hands-on role who will be responsible to build, manage, and automate our Cloud infrastructure. We need someone who will take a key role in our team and help us build our next generation IoT/Embedded devices.
Responsibilities:
- Own the deployment, automation, management, and maintenance of cloud-based production systems, including operational readiness and production reliability.
- Build, deploy, operate, and scale containerized applications using Docker and Kubernetes across complex cloud environments.
- Ensuring availability, performance, security, and scalability of production systems.
- Management of creation, release, and configuration of production systems.
- Evaluation of new technology alternatives and vendor products.
- Own production incident response and resolution, including P1/P2 incident coordination, root-cause analysis, corrective actions, and post-incident reviews.
- Pre-production acceptance testing for quality assurance.
- Design and enforce cloud security controls and IAM policies by applying least-privilege access, secrets management, audit logging, and cloud security best practices.
- Providing recommendations for architecture and process improvements.
- Define and operate comprehensive monitoring, logging, alerting, and observability systems to enable proactive detection and rapid resolution of production issues.
- Design, maintain, and manage automation tools and Infrastructure as Code for repeatable provisioning, configuration, and operational processes.
- Own database backup, restore, replication, and disaster-recovery procedures, including periodic recovery testing and cross-region continuity planning.
- Design and implement testing infrastructures and write code to validate a large scale application deployment, partner with software developers to create/code test tools and implement automation concurrent with product development.
- Collaborate with Software Development teams to maintain consistency and testability across software, databases, migrations, deployment pipelines, and maintenance routines.
- Collaborate with Software Quality Assurance, System Engineering and Architecture, Program Management, and Software Development teams
Required Skills:
- Over 3+ years of experience in provisioning, operations, and management of cloud environments. Experience in creation of highly automated infrastructures.
- Experience in code development in at least one high-level programming language.
- Experience with multi-tier architectures: Load balancers, Caching, Web servers, application servers, databases, and networking.
- Expert level knowledge of cloud platforms (such as AWS, Azure, or Google Cloud).
- Hands on experience in developing Continuous Integration / Continuous Delivery process and expert at implementing pipelines in Jenkins for respective deployment environments.
- Hands on experience with automation / CI-CD tools like Jenkins.
- Hands-on experience designing and operating end-to-end CI/CD pipelines for Java applications using Jenkins, Ansible, Git, Maven, and JFrog, including automated build, test, security checks, release, rollback, and environment promotion.
- Experience in branching, staging and maintaining the version across the environments working on Software Configuration Management (SCM) GIT and GitHub/Bitbucket.
- Hands on experience with microservices integration with the build and deployment pipeline.
- Strong hands-on expertise with Docker and Kubernetes, including container image creation, orchestration, deployment, scaling, networking, security, and troubleshooting.
- Expert knowledge of Building and managing containerized applications, creating Docker images, optimizing resource utilization, and troubleshooting containerization issues.
- Expert knowledge of deploying and managing Kubernetes clusters, configuring and scaling applications using Kubernetes, managing Helm charts, and ensuring high availability and performance.
- Hands-on experience with AWS services including EC2, Auto Scaling, Elastic Load Balancing, S3, EBS, IAM, Route 53, VPC, RDS, CloudWatch, Lambda, EventBridge, AWS Config, and CloudTrail, with strong knowledge of IAM roles, policies, and least-privilege access.
- Hands on experience on various DevOps tools like GIT, Terraform, Jenkins, Ansible, Maven, Prometheus and Grafana
- Hands on experience in Linux and scripting and programming languages (such as Python, Bash, or PowerShell)
- Setup static websites using AWS S3, CloudFront and Route53.
- Hands-on experience with database administration, automated backups, point-in-time recovery, replication, failover, restore validation, and disaster-recovery planning.
- Hands-on experience implementing proactive monitoring, centralized logging, actionable alerting, dashboards, and on-call observability using CloudWatch, Prometheus, Grafana, Elasticsearch, and Kibana.
- Hands-on experience of Continuous monitoring, logging, and analytics to identify performance bottlenecks, security vulnerabilities, and areas for improvement, implementing monitoring and logging solutions (e.g., Prometheus, ELK stack) to enable proactive issue detection and resolution.
- Hands-on experience provisioning new services and environments through Infrastructure as Code using Terraform, with reusable modules, version control, peer review, and controlled state management.
- Update/Modify current infrastructure using Terraform.
- Automate repetitive task/tool using automation playbook (Python Boto3/Ansible/ Terraform)
- Experience with securing cloud infrastructure and meet compliance requirements.
- Experience designing, documenting, implementing, and periodically testing disaster-recovery and business-continuity procedures across regions.
- Demonstrated ownership of high-severity production incidents (P1/P2), including incident command, stakeholder communication, root-cause analysis, postmortems, and follow-through on preventive actions.
- Expert level knowledge of operating system administration.
- Strong networking knowledge.
Preferred Skills:
- Experience operating MQTT brokers and MQTT-based messaging infrastructure, including authentication, authorization, TLS, topic design, monitoring, and troubleshooting.
- Experience with IoT cloud and edge infrastructure, device-to-cloud connectivity, embedded gateways, telemetry pipelines, and large-scale fleet operations.
- Good oral and written communication skills in English
- Must be able to Interface with multiple cross functional teams.