Lead Site Reliability Engineer
ประกาศจากแหล่งภายนอกคุณสมัครได้โดยตรง — เราจะพาคุณไปยังหน้าสมัครงานของบริษัท ไม่ต้องสมัครสมาชิก ไม่มีคนกลาง ไม่ต้องล็อกอิน ThaiJobz
เกี่ยวกับตำแหน่ง
As a Lead Site Reliability Engineer, you'll blend hands-on engineering with team leadership, focusing on platform strategy, reliability, and automation—including AI-assisted operations. This role is ideal for a senior SRE who wants to mentor and lead while solving complex technical challenges.
รายละเอียดงาน
About the Company
At ZILO™, we're redefining what’s possible in technology. ZILO™ is the UK-based FinTech specialising in global asset and wealth management software, designed to scale and transform businesses of all types using our own developed AI Technology. Our mission is to digitalise the future of the global asset management industry. We are a team of experts with decades of combined experience at leading firms globally, who thrive in fast-paced environments and want to shape the future of technology. Every individual plays a key role in driving progress and making a real impact. We continuously strive to innovate and improve.
Why work with us? At ZILO™, you'll be part of a dynamic and inclusive environment where creativity thrives. We offer the opportunity to work on cutting-edge technology, collaborate with talented individuals, and contribute to projects that have a real-world impact. We value continuous learning, personal growth, and providing our team with the resources they need to succeed.
Ready to shape the future? Let’s talk.
Role Details
As Lead Site Reliability Engineer, you'll provide technical and people leadership for the SRE team. You'll be responsible for the team's day-to-day operation, technical direction and engineering delivery, whilst working closely with the Director of Global Cloud Infrastructure to execute the wider platform strategy.
Reporting to the Director of Global Cloud Infrastructure, you'll take ownership of the day-to-day leadership and operation of the Site Reliability Engineering team. Working in close partnership with the Director, you'll translate strategic objectives into operational delivery, ensuring the team consistently delivers a resilient, secure and highly available cloud platform.
You'll work closely with Platform Engineering, Software Engineering, Product, Security and Client Operations teams to improve platform reliability through automation, engineering excellence and operational best practices.
You'll also play a key role in shaping the future of AI-assisted operations at ZILO, identifying opportunities to use AI to improve reliability engineering, incident response, automation and engineering productivity.
This is a hands-on engineering leadership role where approximately 70% of your time will be spent working alongside the team as a Site Reliability Engineer, designing solutions, improving automation and supporting production platforms. The remaining 30% will focus on leading the team, including coaching and mentoring engineers, managing performance, workload planning, holidays, on-call rotas and career development. You'll be expected to lead from the front, setting the technical standard through your own engineering contributions.
This role is not suited to candidates looking to move away from hands-on engineering into full-time management. We're looking for someone who enjoys leading people whilst remaining an active Site Reliability Engineer, spending the majority of their time solving technical challenges alongside the team.
Key Responsibilities
Technical Leadership
- Lead the Site Reliability Engineering team by example, remaining hands-on with the technology and setting the standard for engineering excellence.
- Provide technical guidance and mentorship to SREs, fostering a culture of collaboration, continuous learning and operational excellence.
- Work closely with Platform Engineering and Software Engineering teams to improve the reliability, scalability and operability of our services.
- Champion SRE principles and best practices across the engineering organisation.
- Encourage the adoption of AI-assisted engineering practices, enabling the team to deliver more effectively while maintaining high standards of quality, security and reliability
People Leadership
- Lead, coach and develop a high-performing team of Site Reliability Engineers.
- Conduct regular one-to-one meetings, performance reviews and career development discussions.
- Manage workload, priorities and sprint commitments across the team.
- Set clear objectives and support engineers in achieving individual and team goals.
- Manage holidays, leave requests, on-call rotas and resource planning to ensure effective operational coverage.
- Recruit, onboard and mentor new team members as the team grows.
- Foster a collaborative, accountable and high-performing engineering culture.
Production Operations
- Own the reliability, availability and performance of ZILO's production platforms.
- Work alongside Platform Engineering and Software Engineering teams to support, maintain and patch production environments.
- Define, measure and continually improve Service Level Indicators (SLIs), Service Level Objectives (SLOs) and error budgets.
- Drive continuous improvements to system resilience, fault tolerance and operational readiness.
- Lead platform capacity planning and performance optimisation activities.
Incident Management
- Lead the technical response to high-severity production incidents, providing calm and effective technical leadership during major outages.
- Coordinate post-incident reviews and Root Cause Analysis (RCA), ensuring corrective and preventative actions are identified, prioritised and delivered.
- Drive continuous improvement of incident management processes, reducing Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).
Automation & Platform Engineering
- Eliminate operational toil through automation and engineering-led improvements.
- Design, develop and maintain tooling that improves engineering productivity and operational efficiency.
- Identify opportunities to simplify operational processes through automation and self-service capabilities.
- Work with Platform Engineering to ensure the platform remains scalable, resilient and capable of supporting future business growth.
- Identify opportunities to leverage AI to improve operational efficiency, automate repetitive tasks and enhance engineering productivity.
- Develop and adopt AI-assisted tooling to support troubleshoot
สวัสดิการที่ได้รับ
คุณสมบัติผู้สมัคร
- ประสบการณ์
- มากกว่า 10 ปี
- การศึกษา
- ไม่ระบุ
