Back to Neptunez Singapore Pte. Ltd. jobs
N
Site Reliability Engineer (Dynatrace, GenAI, Prometheus, L2 Support, Splunk)
Islandwide, Singapore
ContractInformation TechnologyJob Description
Responsibilities
- Design, implement, and continuously improve Site Reliability Engineering (SRE) practices to ensure highly available, scalable, and resilient production systems.
- Provide L2 production support by troubleshooting application, infrastructure, and platform issues while ensuring compliance with SLA and SLO commitments.
- Monitor, maintain, and optimize Microsoft Azure environments, including Azure Kubernetes Service (AKS), Azure Monitor, and Log Analytics.
- Support and troubleshoot Java and Spring Boot applications by analyzing logs, JVM performance, REST API issues, configuration problems, and application performance bottlenecks.
- Build, configure, and maintain observability solutions using AppDynamics, Dynatrace, Splunk, Grafana, and Prometheus to improve system visibility and proactive monitoring.
- Develop, enhance, and maintain CI/CD pipelines using Jenkins, Bitbucket, and Azure DevOps to support reliable and automated application deployments.
- Deploy, manage, and troubleshoot containerized applications running on Kubernetes clusters.
- Participate in incident management, on-call rotations, root cause analysis (RCA), and post-incident reviews to improve service reliability and reduce recurring issues.
- Collaborate with development, infrastructure, and DevOps teams to improve application performance, deployment processes, and operational stability.
- Utilize GenAI capabilities to improve operational efficiency through log analysis, incident summarization, automation, and troubleshooting assistance.
- Create and maintain operational documentation, monitoring dashboards, runbooks, and standard operating procedures.
- Manage incidents, change requests, and operational activities using Jira while ensuring timely resolution and effective communication with stakeholders.
Requirements
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
- 8+ years of experience in Site Reliability Engineering (SRE), Production Support, or DevOps environments.
- Strong hands-on experience with Microsoft Azure, including AKS, Azure Monitor, and Log Analytics.
- Experience supporting Java and Spring Boot applications in production environments.
- Experience in Java application troubleshooting, application logs.
- Hand on experience in Microservices, RestAPI
- Hands-on experience with observability and monitoring tools including AppDynamics, Dynatrace, Kibana , Splunk, Grafana, and Prometheus.
- Experience building and maintaining CI/CD pipelines using Jenkins, Bitbucket, and Azure DevOps.
- Hands on experience in managing Kubernetes.
- Experience providing L2 production support, incident management, root cause analysis, and problem management.
- Experience in Jira and change management.
- Exposure to GenAI tools and AI-driven operational automation is an added advantage.
- Strong analytical, troubleshooting, communication, and stakeholder management skills.
- Ability to work in a fast-paced production environment and participate in on-call support rotations.
About Neptunez Singapore Pte. Ltd.
First seen: July 21, 2026
Last updated: July 26, 2026