{
  "contents": [
    {
      "role": "user",
      "parts": [
        {
          "text": "Attachment 1 is a complete English sung advertorial story-song (the inspiration; its words sit tightly on the groove). The other attachments are complete Ukrainian songs telling the same story for a neck cream. Listen to every attachment end to end. For EACH Ukrainian attachment answer with timed evidence (attachment, local seconds, quoted words): (1) inSyncWithBeat: is every phrase in sync with the beat — do the stressed syllables land on the drum beats/eighths, do phrases enter on time? Verdict (tight / mostly tight / loose) and the passages where it drifts early or late or smears a word across beats. (2) songOrNarration: does it sound like a SONG (melody, form, returning material) or like NARRATION set to music? (3) breathes: are consecutive thoughts separated by audible rests of the voice (count clear rests per minute, name three)? (4) sungThroughout: any passage that slips into spoken recitation or patter (local seconds)? (5) codaIntelligible: in the final product/order block list every fact you can make out (brand name, order form with name and phone, inspection at the post office, payment after inspection, a 60-day return, a link below) and whether each has time to land. (6) hookReturns: the returning line and each return time. (7) defects[{localSeconds,words,problem,severity}], wordsMisheardOrUnclear[]. Then across all Ukrainian attachments: whichIsCrisperOnTheBeat (attachment number, why), whichBreathesBest, whichIsMostSong, rankingInSync (best first), rankingOverall (best first) with one sentence each. Return ONLY JSON (no prose outside JSON), <=1900 words, English. Required keys: mediaAccess(boolean), audioAccess(boolean), attachmentsHeard[{attachment,firstWordsHeard,lastWordsHeard,durationEstimateSeconds}], then the analysis keys listed. Quote only words you actually hear; mark uncertain words with (?). Use attachment number plus LOCAL seconds of that attachment. Do not assume any language or attachment is better a priori; do not reward louder mastering, more notes or a softer timbre. Identify concrete mechanisms (where the stressed syllables fall against the kick/snare, early or late phrase entries, words smeared across beats), not taste."
        },
        {
          "text": "Attachment 1: Attachment 1; complete English original song, 194.5s, mono proxy with fixed gain; local time starts at 0."
        },
        {
          "inlineData": {
            "mimeType": "audio/mpeg",
            "data": "[base64 of frozen bytes; see inputs.json]"
          }
        },
        {
          "text": "Attachment 2: Attachment 2; complete Ukrainian song, 228.0s, 40kbps mono proxy with fixed gain; local time starts at 0."
        },
        {
          "inlineData": {
            "mimeType": "audio/mpeg",
            "data": "[base64 of frozen bytes; see inputs.json]"
          }
        },
        {
          "text": "Attachment 3: Attachment 3; complete Ukrainian song, 233.2s, 40kbps mono proxy with fixed gain; local time starts at 0."
        },
        {
          "inlineData": {
            "mimeType": "audio/mpeg",
            "data": "[base64 of frozen bytes; see inputs.json]"
          }
        },
        {
          "text": "Attachment 4: Attachment 4; complete Ukrainian song, 197.1s, 48kbps mono proxy with fixed gain; local time starts at 0."
        },
        {
          "inlineData": {
            "mimeType": "audio/mpeg",
            "data": "[base64 of frozen bytes; see inputs.json]"
          }
        },
        {
          "text": "Attachment 5: Attachment 5; complete Ukrainian song, 190.3s, 48kbps mono proxy with fixed gain; local time starts at 0."
        },
        {
          "inlineData": {
            "mimeType": "audio/mpeg",
            "data": "[base64 of frozen bytes; see inputs.json]"
          }
        }
      ]
    }
  ],
  "stream": true,
  "generationConfig": {
    "thinkingConfig": {
      "includeThoughts": false,
      "thinkingLevel": "high"
    }
  }
}
