DevOps / SRE Engineer - AI Platform
ประกาศจากแหล่งภายนอกMakro PRO
เทคโนโลยี
ทำงานที่ออฟฟิศลงประกาศ 67 วันที่แล้ว
สมัครที่เว็บไซต์บริษัท
คุณสมัครได้โดยตรง — เราจะพาคุณไปยังหน้าสมัครงานของบริษัท ไม่ต้องสมัครสมาชิก ไม่มีคนกลาง ไม่ต้องล็อกอิน ThaiJobz
รายละเอียด
เงินเดือนตามตกลง
ประเภทการจ้าง
เต็มเวลา
รูปแบบ
ทำงานที่ออฟฟิศ
รายละเอียดงาน
About the Role
The DevOps / SRE Engineer owns the operational substrate of an AI-native retail decisioning platform — infrastructure, CI / CD, observability, cost meter, and incident response for a system that runs production agents taking real business actions. The role builds on the enterprise Terraform standard, CI / CD spine, and FinOps tagging policy rather than reinventing parallel infrastructure.
Organization: Makro PRO. Work arrangement: onsite. Remote candidates outside of Thailand are welcome to apply.
Responsibilities
- Adopt the enterprise Terraform standard and module library for all platform infrastructure; author platform-specific modules where needed (agent runtime, vector DB, knowledge graph); run drift detection weekly.
- Build platform-specific CI / CD pipelines on the enterprise spine — service deploys, agent deploys, eval-gate enforcement; integrate eval gates so no agent reaches production without eval pass.
- Operate rollback orchestration with sub-15-minute recovery; run quarterly game days.
- Own the platform observability stack — OpenTelemetry, Langfuse for LLM traces, and custom dashboards for per-agent cost.
- Implement the per-agent cost meter end-to-end — token counts, vector queries, model inference, downstream LLM Gateway costs; surface cost data to the enterprise GenAI cost dashboard.
- Stand up the platform on-call rotation; author runbooks for every production agent and service; lead incident response with measurable corrective actions.
- Implement platform cost-tagging policy consistent with the enterprise standard (team, domain, environment, project, agent, suite, persona); report monthly to Cost Review.
- Drive cost optimisation — right-sizing, caching, model routing decisions, reserved compute.
Qualifications
- Bachelor's or Master's degree in Computer Science, Engineering, or a related discipline.
- 5+ years SRE / DevOps with production ownership.
- Terraform at scale — modules, state, drift, environment promotion.
- CI / CD for data + ML / AI services (GitLab CI / CD or comparable).
- Cloud platform experience (Azure preferred; AWS / GCP transferable).
- Observability experience — OpenTelemetry, Langfuse (or comparable LLM traces), custom dashboards.
- FinOps experience — tagging policies, attribution, optimisation.
- Incident response experience — on-call, post-mortems, runbook authorship.
Preferred Qualifications
- AI / agent platform SRE experience; cost-meter / chargeback systems built or operated.
- Multi-cloud production experience; open-source contributions to IaC / observability tooling.
- AI / ML / agent system observability instrumentation (LLM cost, agent cost, eval scores).
- Vendor certifications such as HashiCorp Terraform Associate / Professional, Azure Solutions Architect Associate, or Databricks Data Engineer Professional.
คุณสมบัติผู้สมัคร
- ประสบการณ์
- 6-10 ปี
- การศึกษา
- ไม่ระบุ
ใบรับรอง / ทักษะเพิ่มเติม
TerraformCI/CDCloud PlatformObservabilityFinOpsIncident ResponseAIMachine LearningCost OptimizationOpenTelemetryLangfuseDashboardsRunbooksProduction OwnershipData Services
