AI Quality Analyst (Personalization) - Japanese
Remote contractor role focused on evaluating Gemini personalization quality using personal Google account data, multi-turn prompts, side-by-side model comparisons, and detailed written rationales.
About this opportunity
Turing is seeking an AI Quality Analyst to evaluate a Gemini personalization feature. The role involves designing multi-turn prompts based on personal experiences, assessing how effectively the model uses information from prior Gemini conversations, Gmail, Google Search, and YouTube activity, and evaluating personalized responses for grounding, integration, helpfulness, naturalness, and overall quality.
What the listing describes
- Design and execute multi-turn conversational prompts, typically 1–5 turns, that require the AI to use personal information and experiences
- Evaluate whether personalization is appropriately applied based on the intent of the starting prompt
- Analyze model responses for grounding issues, unsupported claims, flawed inferences, and hallucinations
- Assess how naturally personal data is integrated into responses without robotic or excessive narration
- Compare and rank two model responses side-by-side based on helpfulness, usability, and overall quality
- Write clear and defensible rationales for model comparisons with references to specific conversation turns
- Extract and verify debug information to confirm that chat summaries and personal data sources were used correctly
- Provide constructive feedback and detailed annotations
- Maintain strict data hygiene by deleting evaluation conversations to prevent them from affecting future chat history
- Collaborate with team members while maintaining high evaluation standards
Requirements shown on the source
- High proficiency in reading and writing Japanese
- Willingness to use a primary personal Google account rather than a testing account
- Willingness to enable personal data sources for evaluation purposes
- Strong analytical ability to assess nuanced and ambiguous AI responses
- Ability to design creative multi-turn prompts based on personal context
- Understanding of personalization concepts, including incorrect personalization, poor inferences, and forced connections
- Strong attention to detail when reviewing side-by-side model responses
- Excellent written communication skills
- Ability to write clear, concise, and structured evaluation rationales
- Ability to provide constructive feedback and detailed annotations
- Strong communication and collaboration skills
- Ability to work independently in a remote environment
- Desktop or laptop with a reliable internet connection
- BS/BA degree or equivalent experience in a relevant analytical field
- Available for at least 4 hours per day and up to 20 hours per week
- Able to provide 4 hours of overlap with PST working hours
Relevant backgrounds mentioned
- Experience in data annotation
- Experience in AI quality evaluation
- Experience in content moderation
- Experience in a related analytical or evaluation role
- Background in Policy, Law, Ethics, Linguistics, Journalism, Computer Science, or a related field
Details to verify before applying
- Country eligibility
- Not explicitly stated
- Compensation
- Listed compensation: USD 15 per hour.
- Last checked
Source and current details
Our summary is not the full job description. Review the external listing for the latest requirements, terms and application process.
Open the source listing on Turing