Remote AI Quality Analyst (Thai)
ประกาศจากแหล่งภายนอกคุณสมัครได้โดยตรง — เราจะพาคุณไปยังหน้าสมัครงานของบริษัท ไม่ต้องสมัครสมาชิก ไม่มีคนกลาง ไม่ต้องล็อกอิน ThaiJobz
รายละเอียดงาน
About the Company
Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems. Turing accelerates frontier research with high-quality data, advanced training pipelines, and AI researchers who specialize in coding, reasoning, STEM, multilinguality, multimodality, and agents; and applies that expertise to help enterprises transform AI from proof of concept into reliable, impactful systems.
About the Role
As an AI Quality Analyst, you will evaluate a new personalization feature for Gemini by assessing how well the model uses information from past Gemini conversations, Gmail, Google Search, and YouTube activity to make responses more relevant and helpful. The role combines creative prompt design and analytical evaluation of personalized model responses across dimensions like Grounding, Integration, and Helpfulness.
Responsibilities
- Design and execute multi-turn conversational prompts (typically 1–5 turns) that require the AI to utilize your personal information and experiences.
- Evaluate model responses based on intent from the starting prompt and verify whether personalization was appropriately applied.
- Analyze responses for Grounding issues to ensure claims about you are supported and not the result of flawed inferences or hallucinations.
- Assess Integration quality to ensure personal data is woven naturally into responses without robotic overnarrating.
- Rigorously evaluate and stack-rank two model responses side-by-side (SxS) to determine which is more helpful, usable, and enjoyable.
- Write clear, defensible rationales for comparisons, explicitly referencing specific turn numbers where issues or positive aspects occurred.
- Extract and verify "Debug Info" from the model to confirm chat summaries and data sources were properly utilized.
- Maintain strict data hygiene by deleting evaluation conversations to prevent them from polluting future chat history.
Qualifications
- Thai proficiency: ability to read and write in Thai with a high degree of competence, as Thai is the focus language for this project.
- Willingness to use your primary personal Google account (not a testing account) and enable personal data sources for genuine assessment.
- Full-time availability in your local time zone; the team staffs a global, 24-hour operations schedule.
- Schedule requirements: at least 4 hours per day and minimum 30 hours per week, with 4 hours overlap with PST (options: 30 hrs/week or 40 hrs/week).
- Exceptional analytical thinking and meticulous attention to detail to evaluate nuanced and ambiguous AI responses.
- Experience in creative prompt engineering: designing multi-turn starting prompts based on personal context.
- Strong evaluation acumen: ability to identify incorrect personalization, poor inferences, and forced connections.
- Excellent written communication: ability to write clear, concise, structured rationales and detailed annotations referencing specific turns.
- Self-motivated and able to work independently in a remote setting; excellent communication and collaboration skills.
- Technical setup: desktop/laptop with a good internet connection.
- Education: BS/BA degree or equivalent experience in a relevant field (e.g., Policy, Law, Ethics, Linguistics, Journalism, Computer Science, or related analytical field).
- Preferred: experience in data annotation, AI quality evaluation, content moderation, or related roles.
Additional Information
- Engagement type: Contractor.
- Engagement length: 3 months.
- After applying, applicants will receive an email with a login link to access the portal and complete their profile.
คุณสมบัติผู้สมัคร
- ประสบการณ์
- 1-2 ปี
- การศึกษา
- ไม่ระบุ
