Media output — media_output
Somewhere the house can hand a stream of audio to and have it come out in a room: the Pi behind a television, a Juno speaker with an amplifier in it, a player belonging to a media server.
The mirror of media_player, and the two are not the same thing said twice. A media player
originates content — it is a source, and a room's listen names one. This is where sound ends
up. It chooses nothing, browses nothing and knows nothing about what it is playing; somebody
upstream decided, and this is the last hop.
Why a separate proxy rather than capabilities on the device that has one. Every output in a
house belongs to something that is already a device for other reasons — the Pi is a navigator,
the Juno speaker is a voice_speaker, a soundbar is a tv's neighbour. Putting has_mixing on
each of those means the same four capabilities written five times and drifting apart, and it means
a room looking for somewhere to put a voice reply has to know all five contracts. One proxy is one
question: which bindings in this room can be given audio, and what happens to what is already
playing when they are.
It carries no audio. Nothing here moves a byte. The same reasoning as voice_speaker's
set_route: whether audio reaches an output as PCM over a socket, as a URL it fetches, or as
something not invented yet is the switching layer's business, and a contract that named a
transport would have to change every time that did. What this contract does say is where to
reach it — endpoint, below, which the driver fills in because the thing serving the audio is the
only party that knows how it is reachable.
Voice, and what happens to the music
The case this exists for is a reply arriving while something is already playing. Three things can happen, and which of them is possible is a fact about the hardware:
- Switching — one audio slot. The media has to stop for the voice to be heard, and start again
afterwards. Anything with
has_mixing = false. - Mixing — two slots. Voice can sound over the media, so the media can either duck under it or
be paused anyway.
has_mixing = true, and which of the two happens is a decision somebody makes per output, not something the device decides.
Ducking is not always the kinder option, which is why it is a choice rather than a default: under a film at volume, a ducked track is still loud enough to bury a short answer, and a pause that resumes cleanly is less annoying than an answer nobody could make out.
Capabilities
Declared in the driver manifest under [[proxy]] capabilities. Anything not declared takes the default below.
| Capability | Type | Default | Meaning |
|---|---|---|---|
has_media | bool | true | A media stream can be played here |
has_mixing | bool | false | Voice and media can sound at once |
has_mute | bool | false | |
has_pause | bool | false | Media playing here can be paused and resumed in place |
has_voice | bool | true | Voice replies and announcements can be played here |
has_volume | bool | false | |
services | string | "url" | Comma-separated protocols this output can host |
volume_max | u32 | 100 |
Commands
pause
Only present when
has_pauseis declared true.
Stop playing and keep the place. See has_pause for why not every output has this: a stream is
not a file, and an internet radio station that is paused for four seconds and resumed has moved on
four seconds.
No parameters.
play
Only present when
has_mediais declared true.
Fetch this and play it. The whole of what an output needs to be told.
A URL, not audio. The controller never carries a byte of it — a continuous stream through the one process the house depends on is the thing every other part of this design goes out of its way to avoid, and it would put the controller between a media server and a speaker for no benefit. The output fetches for itself, which also means a source too big or too fast for the controller's link is not thereby too big for the house.
at is what makes several rooms one room. An empty at means "as soon as you can", which is
right for one output. Given a moment, the output prepares and waits for it — so a controller
handing three outputs the same URL and the same instant gets three rooms that start together
instead of three rooms a second apart, which in an open-plan floor is worse than one room.
RFC 3339, with milliseconds, in UTC. A wall clock rather than a countdown because a countdown is wrong by however long the request took, and that is exactly the error being removed. Everything in a house runs NTP; an output whose clock is wrong will start at the wrong time and is worth fixing rather than working around.
The metadata is for whatever is drawing a screen in that room. Optional in the honest sense — a stream that has no title has none, and inventing "Unknown Artist" is a decision for whoever draws it rather than for the thing playing it.
| Parameter | Type | Notes |
|---|---|---|
artist | string (optional) | |
artwork | string (optional) | URL of cover art |
at | string (optional) | RFC 3339 instant to start at. Empty means now. |
title | string (optional) | |
url | string |
resume
Only present when
has_pauseis declared true.
Carry on from where pause stopped.
No parameters.
set_mute
Only present when
has_muteis declared true.
| Parameter | Type | Notes |
|---|---|---|
mute | bool |
set_services
Connect this endpoint to the selected media coordinator and advertise it under this name.
configuration is an opaque JSON object owned by the named service. For Music Assistant it
contains the Sendspin address and the stable player id Core assigned this output; it never contains
the Music Assistant access token.
The whole list every time, never a delta. The same bargain HostCall::Connections makes and for
the same reason: after a restart, a reconnect or a controller that was replaced, one message says
what should be true and there is no resync path to get wrong. An empty list means host nothing,
which is how a room is removed from the service without affecting independent voice audio.
name is the room label shown for this output inside the service, rather than the box hostname.
Empty leaves naming to the endpoint.
Anything in services that this output did not declare it can host is ignored rather than
refused: a controller that has learned about a protocol before the box has is a controller in the
middle of an upgrade, and it should not take the working half down.
| Parameter | Type | Notes |
|---|---|---|
configuration | string (optional) | Service-owned JSON configuration; never credentials |
name | string (optional) | |
services | string | Comma separated, from this output's own services |
set_volume
Only present when
has_volumeis declared true.
| Parameter | Type | Notes |
|---|---|---|
level | u32 ≥0 |
stop
Stop whatever is playing here. Does not turn anything off — an output is not a device.
No parameters.
Notifications
endpoint_changed
Where audio for this output is to be sent, as the thing serving it understands the question.
Reported rather than derived, and for the reason voice_speaker's listening_changed reports one:
whoever holds the audio is the only party that knows where it is reachable, and anything that
guessed would be wrong the first time a house ran two of them. Empty means it cannot currently be
reached, which is a different statement from being offline.
| Parameter | Type | Notes |
|---|---|---|
endpoint | string |
mute_changed
Only present when
has_muteis declared true.
| Parameter | Type | Notes |
|---|---|---|
mute | bool |
now_playing_changed
Only present when
has_mediais declared true.
What this output is playing now, as it turned out — not as it was asked.
Reported rather than assumed from the command, and the difference is the whole point: a URL that redirects, a station that names its own current track, a stream whose duration is only known once the headers arrive. A controller that echoed back what it sent would show the wrong title for every internet radio station in the world.
An empty url means nothing is playing, which is how an output says a track ended or a stream
dropped without anybody having asked it to stop.
| Parameter | Type | Notes |
|---|---|---|
artist | string (optional) | |
artwork | string (optional) | |
duration | u32 (optional) | Seconds. Absent for a live stream, which has none. |
source | string (optional) | url, music_assistant, or another declared service — empty when nothing is playing |
title | string (optional) | |
url | string |
online_changed
| Parameter | Type | Notes |
|---|---|---|
online | bool |
playing_changed
Whether media, voice, or both are coming out of this output right now.
Two booleans rather than one state word because on a mixing output both are true at once, and a
single playing/speaking/idle would have to invent a name for that or lie about it.
| Parameter | Type | Notes |
|---|---|---|
media | bool | |
voice | bool |
position_changed
Only present when
has_mediais declared true.
How far into the track this output is, in seconds.
Sent on the edges that matter — a start, a seek, a pause — rather than every second. A progress bar wants a number and a clock, not a stream of numbers: whoever is drawing one has the wall clock already and can move the bar itself between reports, and a house with six outputs in it should not be sending six messages a second to say that time is passing.
| Parameter | Type | Notes |
|---|---|---|
seconds | u32 |
volume_changed
Only present when
has_volumeis declared true.
| Parameter | Type | Notes |
|---|---|---|
level | u32 |
State
Last-known values core keeps for a binding of this proxy.
| Key | Type | Meaning |
|---|---|---|
artist | string | |
artwork | string | |
duration | u32 | |
endpoint | string | |
media | bool | |
mute | bool | |
online | bool | |
position | u32 | |
source | string | |
title | string | |
url | string | |
voice | bool | |
volume | u32 |