Home / News EN / Apple Watch Series 12 and Ultra 4 Introduce Audio Intelligence with Privacy Focus

Apple Watch Series 12 and Ultra 4 Introduce Audio Intelligence with Privacy Focus

The new Apple Watch Series 12 and Apple Watch Ultra 4 are capable of listening to the surrounding environment throughout the day thanks to the Audio Intelligence feature. This functionality, described in official Apple support documentation, recognizes critical sounds such as alarms or crying babies, identifies musical tracks via integrated Shazam, allows playback of the last 15 seconds of a conversation in text form (Live Rewind), and generates a summary of daily conversations (Siri Recap), operating without the need for manual activation.

Audio Intelligence privacy

Apple has designed this feature around a privacy-centric architecture. The new S11 chip includes a Secure Enclave, a hardware-isolated compartment where audio is transcribed and immediately deleted, both on the watch and the associated iPhone. Only condensed text, devoid of speaker identification, reaches Apple’s private cloud for final summary generation: neither third-party applications, nor the operating system, nor the user themselves, nor Apple can access the original audio, and each function can be disabled individually.

Audio Intelligence privacy: why it matters

The four Audio Intelligence features are available on the Apple Watch Series 12 and Ultra 4, on sale from September 18 starting at €459 and €909 respectively. Each remains deactivated until the user intentionally enables it via iPhone settings or the Watch app.

Sound Recognition identifies sirens, alarms, and doorbells, providing alerts even when the iPhone is not nearby, making it particularly useful for people with hearing difficulties. Music Recognition, based on Shazam, works automatically by displaying the title and artist in the Smart Stack without requiring the widget to be opened. For both functions, the Secure Enclave sends only an acoustic fingerprint—a compressed data point that cannot be reconstructed into the original sound—to servers.

Live Rewind is activated with a double press of the Digital Crown and retrieves the last 15 seconds of speech as text viewable on screen; the fragment automatically disappears approximately 30 seconds after the screen turns off, unless saved in the Siri app. Activation is signaled by an acoustic signal emitted from the speaker, even when volume is silenced, and by a full-screen animation with the active microphone icon. This represents the only aspect where Apple extends privacy considerations to surrounding people.

Siri Recap takes environmental notes throughout the day to generate a high-level summary, which is automatically deleted after 7 days if not saved by the user. The text sent to the iPhone for final synthesis on Private Cloud Compute is designed to be less than half the length of the original transcription, stripped of tone, fillers, and repetitions. The summary intentionally omits attribution of sentences to specific individuals and excludes sensitive categories such as financial data and identity documents, according to technical documentation published by Apple.

What Changes and What Are the Effects

Live Rewind and Siri Recap remain in beta until the end of the year, available only in English and requiring at least an iPhone 16 with Apple Intelligence and the new AI-powered Siri, which is also still under testing. Usage is prohibited for those under 13 years old, and for teenagers between 13 and 17, it requires activation by a parent on a Child Account. In Europe, the launch of AI-powered Siri in Italy is hindered by the Digital Markets Act, the same regulatory constraint that has already delayed other Apple Intelligence features in the European Union.

In 2019, an investigation revealed that thousands of employees were listening to and transcribing Alexa audio clips, including private conversations accidentally captured—a risk previously flagged by the Data Protection Authority for enterprise wearables collecting biometric data at work. Apple has implemented Audio Intelligence on a different architecture: isolated processing in silicon, condensed text before cloud transmission, and no speaker attribution.

The S11 chip must prove its reliability through this architecture, step by step. For corporate purchases, the priority is understanding where the text ends after the microphone turns off, rather than relying solely on the security promises contained in the press release.

Source and further reading on Audio Intelligence privacy: original article.

* Content created with the assistance of artificial intelligence systems.