Remote AI Quality Analyst (Thai)
ประกาศจากแหล่งภายนอกคุณสมัครได้โดยตรง — เราจะพาคุณไปยังหน้าสมัครงานของบริษัท ไม่ต้องสมัครสมาชิก ไม่มีคนกลาง ไม่ต้องล็อกอิน ThaiJobz
รายละเอียดงาน
About the Company
Based in San Francisco, California, Turing is a research accelerator for frontier AI labs and a partner for enterprises deploying advanced AI systems. Turing accelerates research with high-quality data, training pipelines, and AI researchers, and helps enterprises move AI from proof of concept to production.
About the Role
As an AI Quality Analyst, you will evaluate a personalization feature for Gemini by assessing how the model uses information from past Gemini conversations, Gmail, Google Search, and YouTube activity to make responses more relevant and helpful. The role combines creative prompt design and analytical evaluation of personalized model outputs.
Responsibilities
- Design and execute multi-turn conversational prompts (typically 1–5 turns) that require the AI to utilize your personal information and experiences.
- Evaluate model responses against the intent in the starting prompt and check whether personalization was appropriately applied.
- Analyze responses for grounding issues, flawed inferences, hallucinations, and incorrect personalization.
- Assess integration quality to ensure personal data is woven naturally without overnarrating.
- Compare and stack-rank two model responses side-by-side (SxS) to determine which is more helpful, usable, and enjoyable.
- Write clear, defensible rationales for comparisons, explicitly referencing specific turn numbers.
- Extract and verify debug info to confirm chat summaries and data sources were used correctly.
- Maintain strict data hygiene by deleting evaluation conversations to avoid polluting future chat history.
Qualifications
- Thai proficiency: ability to read and write in Thai at a high level (Thai is the project focus).
- Willingness to use your primary personal Google account and enable personal data sources for genuine assessment.
- Full-time availability in your local time zone with schedule flexibility; part of a global 24-hour operations team.
- Exceptional analytical thinking and the ability to evaluate nuanced and ambiguous AI responses.
- Experience designing creative, multi-turn starting prompts (prompt engineering).
- Strong evaluation acumen: identify incorrect personalization, poor inferences, and forced connections.
- Meticulous attention to detail and ability to spot subtle differences in naturalness and overnarrating.
- Excellent written communication: ability to write clear, concise, structured rationales and annotations.
- Self-motivated and able to work independently in a remote setting.
- BS/BA degree or equivalent experience in Policy, Law, Ethics, Linguistics, Journalism, Computer Science, or related analytical field preferred.
- Experience in data annotation, AI quality evaluation, content moderation, or related roles is strongly preferred.
- Desktop/laptop with a good internet connection and appropriate technical setup.
Benefits
- Engagement type: Contractor.
- Engagement length: 3 months.
- Time commitment options: 30 hours/week (minimum) or 40 hours/week; at least 4 hours per day and minimum 30 hours per week required.
- Must provide 4 hours of overlap with PST.
Additional Information
- After applying, you will receive an email with a login link to access the portal and complete your profile.
- Maintain strict data hygiene: delete evaluation conversations after completion to prevent future chat pollution.
คุณสมบัติผู้สมัคร
- ประสบการณ์
- 1-2 ปี
- การศึกษา
- ไม่ระบุ
