Job Overview
SRE in OneDegree is a global team, our mission is to be a business enabler to accelerate the time to release and provide our customers with a reliable and secure digital journey.
We are seeking a skilled and passionate Site Reliability Engineer with a strong technical background and excellent communication skills. This individual will lead the development, construction, and management of reliable and distributed systems that support our business operations.
In this role, you will play a vital part in supporting the following businesses:
- IXT: An insurance core system solution for APAC insurance markets.
- OneDegree HK: A user-friendly digital insurance platform for individuals and businesses in Hong Kong.
- Cymetrics: A cybersecurity platform designed specifically for small and medium enterprises in the APAC region.
Know more about SRE team in OneDegree👉
https://medium.com/onedegree-tech-blog/onedegree-sre-%E5%9C%98%E9%9A%8A%E5%A4%A7%E6%8F%AD%E7%A7%98-385548018fb8
Responsibilities
- Implement and enhance system reliability, availability, scalability, performance, and efficiency by leveraging monitoring, alerting, and automation tools on public cloud platforms like Azure and GCP.
- Participate in capacity planning, analyze software performance, and fine-tune systems to ensure optimal operation.
- Develop and enhance our CI/CD process and toolset to streamline software delivery and deployment.
- Define and monitor key metrics to assess and enhance system reliability.
- Collaborate closely with the engineering team to improve reliability and operational efficiency at every software development life cycle (SDLC) stage.
- Troubleshoot, optimize infrastructure and automate repetitive tasks to increase efficiency and effectiveness.
Requirements
- Proficiency in programming languages such as Bash, Python, or Go.
- Advanced knowledge of monitoring solutions like Prometheus, Grafana, ELK (Elasticsearch, Logstash, Kibana).
- Strong expertise and experience in cloud technologies, specifically Azure and GCP.
- Experience in the complete software development life cycle (SDLC).
- In-depth understanding of network concepts, particularly with a focus on security.
- Hands-on experience implementing CI/CD processes, for example, using GitLab CI.
- Proficiency in automation platforms like Ansible and Terraform.
- Knowledge of orchestration tools like Kubernetes.
- Familiarity with container technologies like Docker.
- Experience with Git source code version control systems.
- Strong problem-solving skills with a systematic approach, effective communication abilities, and a self-driven attitude.
Interview Process
- HR phone interview
- 1st Interview: 1.5~2 hours, 1 hour meet with hiring team + 0.5 hours with HR
- 2nd Interview: 0.5~1 hour, meet with CTO
Compensation commensurate with experience.