Repository navigation
Conversation
Contributor
There was a problem hiding this comment.
🔍 Devin Review: 1 flag
Not posted on this PR by your GitHub settings — view it in Devin Review. (Configure)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
This is the first step toward a shared sans-io core for the voice pipeline that both
livekit/agents(Python) andlivekit/agents-jscould eventually run as a single Rust implementation. Endpointing is the proving ground for three reasons: it's small, the two SDKs already maintain hand-ported copies of it, and its logic was already almost pure.This PR writes the core in pure Python and defines a language-neutral contract that a later Rust or JS core has to pass. It changes no behavior.
What
livekit.agents.core.endpointingis a sans-io reducer.UserSpeechStarted,UserSpeechEnded,AgentSpeechStarted,AgentSpeechEnded,OptionsUpdated.handle(event):MinDelayUpdated,UtteranceEndAdjusted,NonInterruptionOverridden.FixedEndpointingandDynamicEndpointingexpose read-onlymin_delay/max_delay/overlapping.at), never logs.livekit.agents.voice.endpointingkeepsBaseEndpointing,DynamicEndpointingandcreate_endpointingwith identical signatures. These are now thin subclasses of the core: each method call becomes an event, and the returned outputs are logged with the same message, level andextrakeys as before.AudioRecognitionandAgentActivityare untouched.tests/core_fixtures/endpointing.jsonholds 13 scenarios. Each step lists the event, the exact outputs, and the resulting state, and together they cover every branch of the dynamic logic. They were recorded from the pre-refactor implementation.tests/test_core_endpointing.pyreplays them against the core with an absolute float tolerance of 1e-9.Commits
refactor(endpointing): move logic into livekit.agents.core: a pure move with no behavior change.refactor(endpointing): sans-io reducer core: adds the events, outputs andhandle(). The shell translates method calls and logs.test(endpointing): language-neutral core fixtures: adds the fixture contract, its runner, and the stdlib-only import test.Verification
pytest --unit: 4424 passed, 5 skipped. Existing endpointing, agent-session and STT-trace tests pass, and the only edit to them is one appended regression test for explicitinterruption=None.mainand on this branch, capturing log records and state. Both runs produced byte-identical output.make type-checkandruffare clean.Not in this PR
agents/src/voice/turn_config/endpointing.tsis unchanged. A follow-up can runtests/core_fixtures/endpointing.jsonagainst it to check parity. Itsinterruption === falsecheck already matches the core's semantics.