ตำแหน่งงาน
ที่ไหน

18,236 ตำแหน่ง

กลับไปยังหน้าหางาน
อัพโหลด CV ของคุณ

เราจะใช้ AI อ่าน แล้วช่วยหาโอกาสที่ใช่ให้คุณ

คลิกเพื่ออัพโหลด หรือลากไฟล์มาวาง · PDF, DOC, DOCX · ไม่เกิน 10 MB

ttb bank

System Reliability Engineer (Platform)

ประกาศจากแหล่งภายนอก
ttb bank
เทคโนโลยี
ทำงานที่ออฟฟิศลงประกาศ 49 วันที่แล้ว
สมัครที่เว็บไซต์บริษัท

คุณสมัครได้โดยตรง — เราจะพาคุณไปยังหน้าสมัครงานของบริษัท ไม่ต้องสมัครสมาชิก ไม่มีคนกลาง ไม่ต้องล็อกอิน ThaiJobz

รายละเอียด
เงินเดือนตามตกลง
ประเภทการจ้าง
เต็มเวลา
รูปแบบ
ทำงานที่ออฟฟิศ

รายละเอียดงาน

About the Role

System Reliability Engineer (Platform) responsible for operating and improving the platform that runs AKS clusters, API gateways, and AI services.

Responsibilities

  • Monitor, maintain, and optimize Azure Kubernetes Service (AKS) cluster health, availability, and performance.
  • Define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for all platform services.
  • Lead incident response across L1 (automated detection and recovery), L2 (root cause investigation), and L3 (deep engineering resolution); produce post-mortem documentation for P1/P2 incidents.
  • Implement and enforce cluster security policies, resource quotas, and namespace governance.
  • Manage the central API gateway including authentication policies, rate limiting, blue-green traffic routing, and request monitoring; ensure zero-downtime deployments with instant rollback capability.
  • Operate the central abstraction layer for Azure OpenAI and Azure AI Services — model version control, endpoint configuration, usage monitoring, and reliability mechanisms (retry, fallback, failover).
  • Monitor AI service usage and API traffic patterns; detect anomalies, adjust configurations proactively, and maintain SLAs for platform-managed AI services.
  • Build and maintain an observability stack covering API metrics, resource metrics, log standards, and distributed tracing across platform services.
  • Develop dashboards and alerting mechanisms across all support tiers; every alert must include service name, error type, likely cause, and trace ID.
  • Establish and maintain runbooks for L1 triage, L2 escalation, and L3 resolution across all critical platform components.
  • Design, implement, and maintain CI/CD pipelines using Jenkins and Azure DevOps for automated deployment of data and AI services.
  • Enforce code quality, testing standards, and deployment governance through pipeline automation.
  • Manage Kubernetes-based application deployments across UAT and production environments.
  • Monitor and optimize cloud resource utilization and spending across Azure services; conduct capacity planning and implement resource tagging, cost allocation, and rightsizing recommendations.

Qualifications

  • Bachelor's degree or higher in Computer Science, Information Technology, Software Engineering, or a related field.
  • 5+ years of hands-on experience in data engineering, platform engineering, DevOps, or SRE roles.
  • Demonstrated experience operating Kubernetes clusters in production environments, preferably on Azure (AKS).
  • Experience handling platform support across L1 (automated monitoring & recovery), L2 (escalation & triage), and L3 (deep engineering & resolution) tiers.
  • Experience with API gateway management in a microservices environment is advantageous.
  • Experience in a financial services or regulated environment is advantageous.

Skills

  • Required: Python, Bash/shell scripting, Jenkins, Azure DevOps, observability and monitoring (metrics, logs, distributed tracing), Kubernetes (AKS), CI/CD pipeline design and automation.
  • Preferred: Helm & Kubernetes manifests, API gateway platforms (Kong, NGINX, Azure API Management), Azure OpenAI or equivalent AI service operations, Terraform or equivalent IaC, Azure CLI, Databricks workspace administration.

คุณสมบัติผู้สมัคร

ประสบการณ์
6-10 ปี
การศึกษา
ไม่ระบุ
ใบรับรอง / ทักษะเพิ่มเติม
PythonBashShell ScriptingCI/CD ToolingObservabilityHelmKubernetesAPI GatewayAzure OpenAITerraformAzure CLIDatabricksMonitoringLoggingIncident Response

เกี่ยวกับบริษัท