Remote AI Quality Analyst (Thai)
ประกาศจากแหล่งภายนอกคุณสมัครได้โดยตรง — เราจะพาคุณไปยังหน้าสมัครงานของบริษัท ไม่ต้องสมัครสมาชิก ไม่มีคนกลาง ไม่ต้องล็อกอิน ThaiJobz
รายละเอียดงาน
About the Company
Based in San Francisco, California, Turing is a research accelerator for frontier AI labs and a partner for enterprises deploying advanced AI systems. Turing supports frontier research with high-quality data, training pipelines, and AI researchers, and helps enterprises deploy reliable AI that delivers measurable impact.
About the Role
As an AI Quality Analyst, you will evaluate a personalization feature for Gemini, assessing how the model uses information from past Gemini conversations, Gmail, Google Search, and YouTube activity to produce more relevant and helpful responses. The role combines creative prompt design and rigorous analytical evaluation of personalized model outputs.
Responsibilities
- Design and execute multi-turn conversational prompts (typically 1–5 turns) that require the AI to use your personal information and experiences.
- Evaluate model responses against the intent of the starting prompt and check whether personalization was applied appropriately.
- Analyze responses for grounding issues to ensure claims about you are supported by evidence and not flawed inferences or hallucinations.
- Assess integration quality to ensure personal data is woven naturally without robotic overnarrating.
- Rigorously compare and stack-rank two model responses side-by-side (SxS) to determine which is more helpful, usable, and enjoyable.
- Write clear, defensible rationales for comparisons, explicitly referencing specific turn numbers.
- Extract and verify debug info from the model to confirm chat summaries and data sources were used properly.
- Maintain strict data hygiene by deleting evaluation conversations to avoid polluting future chat history.
Qualifications
- Thai proficiency: ability to read and write Thai at a high level of comprehension, as Thai is the focus language for this project.
- Willingness to use your primary personal Google account (not a testing account) and enable personal data sources for genuine assessment.
- Full-time availability in your local time zone with schedule flexibility; part of a global 24-hour operations team.
- At least 4 hours per day and minimum 30 hours per week, with 4 hours overlap with PST (options: 30 hrs/week or 40 hrs/week).
- Exceptional analytical thinking and the ability to evaluate nuanced and ambiguous AI responses, specifically personalization quality.
- Experience in creative prompt engineering and designing multi-turn starting prompts based on personal context.
- Strong evaluation acumen: ability to identify incorrect personalization, poor inferences, and forced connections.
- Meticulous attention to detail and ability to spot subtle differences in SxS responses.
- Excellent written communication: write clear, concise, structured rationales and detailed annotations.
- Self-motivated and able to work independently in a remote setting.
- Technical setup: desktop or laptop with a reliable internet connection.
- BS/BA degree or equivalent experience in a relevant field (Policy, Law, Ethics, Linguistics, Journalism, Computer Science, or related analytical field) preferred.
- Experience in data annotation, AI quality evaluation, or content moderation is strongly preferred.
Offer Details
- Engagement type: Contractor.
- Engagement length: 3 months.
- Commitments required: at least 4 hours per day and minimum 30 hours per week, with 4 hours overlap with PST; two options of time commitment: 30 hrs/week or 40 hrs/week.
- After applying, you will receive an email with a login link to access the portal and complete your profile.
คุณสมบัติผู้สมัคร
- ประสบการณ์
- 1-2 ปี
- การศึกษา
- ไม่ระบุ
