Remote AI Quality Analyst (Thai)
ประกาศจากแหล่งภายนอกคุณสมัครได้โดยตรง — เราจะพาคุณไปยังหน้าสมัครงานของบริษัท ไม่ต้องสมัครสมาชิก ไม่มีคนกลาง ไม่ต้องล็อกอิน ThaiJobz
รายละเอียดงาน
About the Company
Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing supports customers by accelerating frontier research with high-quality data, advanced training pipelines, and top AI researchers, and by helping enterprises transform AI from proof of concept into reliable, impactful systems.
About the Role
As an AI Quality Analyst, you will evaluate a new personalization feature for Gemini, assessing how well the model uses information from past Gemini conversations, Gmail, Google Search, and YouTube activity to make responses more relevant and helpful. The role combines creativity and analytical rigor: you will design prompts based on personal experiences and evaluate personalization dimensions such as Grounding, Integration, and Helpfulness.
Responsibilities
- Design and execute multi-turn conversational prompts (typically 1–5 turns) that require the AI to utilize your personal information and experiences.
- Evaluate model responses based on your intent from the starting prompt and check if personalization was appropriately applied.
- Analyze responses for Grounding issues, ensuring claims about you are supported by evidence and not flawed inferences or hallucinations.
- Assess Integration quality to ensure personal data is woven naturally into the response without robotic overnarrating.
- Rigorously evaluate and stack-rank two model responses side-by-side (SxS) to determine which is overall more helpful, easy to use, and enjoyable.
- Write clear, defensible rationales for comparisons, explicitly referencing specific turn numbers.
- Extract and verify "Debug Info" from the model to confirm that chat summaries and data sources were properly utilized.
- Maintain strict data hygiene by deleting evaluation conversations to prevent them from polluting your future chat history.
Qualifications
- Thai proficiency: ability to read and write in Thai with a high degree of competence (Thai is the focus language for this project).
- Willingness to use your primary personal Google account (not a testing account) and enable personal data sources for a genuine assessment.
- Schedule flexibility: full-time availability in your local time zone is required; we are staffing a global, 24-hour operations team.
- Commitments required: at least 4 hours per day and minimum 30 hours per week with 4 hours of overlap with PST. Options: 30 hrs/week or 40 hrs/week.
- Exceptional analytical thinking and strong evaluation acumen to assess personalization concepts, identify incorrect personalization, poor inferences, and forced connections.
- Creative prompt engineering: experience designing creative, multi-turn starting prompts based on personal context to thoroughly test model capabilities.
- Meticulous attention to detail and excellent written communication; ability to write clear, concise, and structured rationales referencing specific turns.
- Experience in data annotation, AI quality evaluation, content moderation, or a related role is strongly preferred.
- BS/BA degree or equivalent experience in Policy, Law, Ethics, Linguistics, Journalism, Computer Science, or a related analytical field.
- Technical setup: desktop/laptop with a good internet connection.
Additional Information
- Engagement type: Contractor. Engagement length: 3 months.
- After applying, you will receive an email with a login link. Use that link to access the portal and complete your profile.
คุณสมบัติผู้สมัคร
- ประสบการณ์
- 1-2 ปี
- การศึกษา
- ไม่ระบุ
