Remote AI Quality Analyst (Thai)
ประกาศจากแหล่งภายนอกคุณสมัครได้โดยตรง — เราจะพาคุณไปยังหน้าสมัครงานของบริษัท ไม่ต้องสมัครสมาชิก ไม่มีคนกลาง ไม่ต้องล็อกอิน ThaiJobz
รายละเอียดงาน
About the Company
Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing accelerates frontier research with high-quality data and advanced training pipelines, and helps enterprises transform AI from proof of concept into reliable proprietary intelligence that delivers measurable impact.
About the Role
As an AI Quality Analyst, you will evaluate a new personalization feature for Gemini. You will assess how well the model uses information from past Gemini conversations, Gmail, Google Search, and YouTube activity to make responses more relevant and helpful. The role combines creative prompt design and analytical evaluation of model personalization across dimensions such as Grounding, Integration, and Helpfulness.
Responsibilities
- Design and execute multi-turn conversational prompts (typically 1–5 turns) that require the AI to utilize your personal information and experiences.
- Evaluate model responses against the starting prompt intent and check if personalization was appropriately applied.
- Analyze responses for grounding issues to ensure claims about you are supported and not flawed inferences or hallucinations.
- Assess integration quality to ensure personal data is woven naturally into responses without robotic overnarrating.
- Stack-rank two model responses side-by-side (SxS) to determine which is overall more helpful, easy to use, and enjoyable.
- Write clear, defensible rationales for comparisons, explicitly referencing specific turn numbers.
- Extract and verify debug information to confirm chat summaries and data sources were properly utilized.
- Maintain strict data hygiene by deleting evaluation conversations to prevent polluting future chat history.
Key Qualifications
- Ability to read and write Thai with a high degree of competence; Thai is the focus language for this project.
- Willingness to use your primary personal Google account (not a testing account) and enable personal data sources for genuine assessment.
- Full-time availability in your local time zone; part of a global, 24-hour operations team.
- Commitment of at least 4 hours per day and minimum 30 hours per week with 4 hours overlap with PST; options: 30 hrs/week or 40 hrs/week.
- Exceptional analytical thinking and strong evaluation acumen for assessing personalization quality, incorrect personalization, poor inferences, and forced connections.
- Experience designing creative, multi-turn prompts and experience in data annotation, AI quality evaluation, or content moderation is strongly preferred.
- Meticulous attention to detail and excellent written communication; able to write concise, structured rationales referencing specific turns.
- Self-motivated and able to work independently in a remote setting; desktop/laptop with a good internet connection required.
- BS/BA degree or equivalent experience in a relevant field (Policy, Law, Ethics, Linguistics, Journalism, Computer Science, or related analytical field).
Benefits
- Engagement type: Contractor. Engagement length: 3 months.
- Commitments required: at least 4 hours per day; minimum 30 hours per week with 4 hours overlap with PST (option for 40 hrs/week).
Additional Information
- After applying, you will receive an email with a login link to access the portal and complete your profile.
คุณสมบัติผู้สมัคร
- ประสบการณ์
- 1-2 ปี
- การศึกษา
- ไม่ระบุ
