EXTERNAL LISTING · TURING

AI Quality Analyst (Personalization) - Japanese

Remote contractor role focused on evaluating Gemini personalization quality using personal Google account data, multi-turn prompts, side-by-side model comparisons, and detailed written rationales.

JapaneseAI Quality EvaluationRemote

About this opportunity

Turing is seeking an AI Quality Analyst to evaluate a Gemini personalization feature. The role involves designing multi-turn prompts based on personal experiences, assessing how effectively the model uses information from prior Gemini conversations, Gmail, Google Search, and YouTube activity, and evaluating personalized responses for grounding, integration, helpfulness, naturalness, and overall quality.

What the listing describes

  • Design and execute multi-turn conversational prompts, typically 1–5 turns, that require the AI to use personal information and experiences
  • Evaluate whether personalization is appropriately applied based on the intent of the starting prompt
  • Analyze model responses for grounding issues, unsupported claims, flawed inferences, and hallucinations
  • Assess how naturally personal data is integrated into responses without robotic or excessive narration
  • Compare and rank two model responses side-by-side based on helpfulness, usability, and overall quality
  • Write clear and defensible rationales for model comparisons with references to specific conversation turns
  • Extract and verify debug information to confirm that chat summaries and personal data sources were used correctly
  • Provide constructive feedback and detailed annotations
  • Maintain strict data hygiene by deleting evaluation conversations to prevent them from affecting future chat history
  • Collaborate with team members while maintaining high evaluation standards

Requirements shown on the source

  • High proficiency in reading and writing Japanese
  • Willingness to use a primary personal Google account rather than a testing account
  • Willingness to enable personal data sources for evaluation purposes
  • Strong analytical ability to assess nuanced and ambiguous AI responses
  • Ability to design creative multi-turn prompts based on personal context
  • Understanding of personalization concepts, including incorrect personalization, poor inferences, and forced connections
  • Strong attention to detail when reviewing side-by-side model responses
  • Excellent written communication skills
  • Ability to write clear, concise, and structured evaluation rationales
  • Ability to provide constructive feedback and detailed annotations
  • Strong communication and collaboration skills
  • Ability to work independently in a remote environment
  • Desktop or laptop with a reliable internet connection
  • BS/BA degree or equivalent experience in a relevant analytical field
  • Available for at least 4 hours per day and up to 20 hours per week
  • Able to provide 4 hours of overlap with PST working hours

Relevant backgrounds mentioned

  • Experience in data annotation
  • Experience in AI quality evaluation
  • Experience in content moderation
  • Experience in a related analytical or evaluation role
  • Background in Policy, Law, Ethics, Linguistics, Journalism, Computer Science, or a related field

Details to verify before applying

Country eligibility
Not explicitly stated
Compensation
Listed compensation: USD 15 per hour.
Last checked

Source and current details

Our summary is not the full job description. Review the external listing for the latest requirements, terms and application process.

Open the source listing on Turing