| Age | Commit message (Collapse) | Author | Files | Lines |
|
|
|
start() previously logged only at entry, so a silent throw between
'Scrypted Viewport up' and 'this.running = true' would leave us
guessing. Add explicit end-of-method logs with the post-write state,
plus a try/catch around bootstrap that surfaces the exception
message before re-throwing.
This will tell us whether start() reaches its set this.running=true
on script load (i.e. whether the prototype state-proxy write fires)
or whether bootstrap is silently failing somewhere mid-way.
|
|
|
|
panel binding
The 'Status and Controls' panel in @scrypted/core 0.3.147 was rendering
Unknown despite the previous commit adding StartStop. Reading the
SDK (sdk/src/index.ts:195 + plugins/core/src/script.ts:47) confirmed
how it should work:
- ScryptedDeviceBase installs Object.defineProperty getters/setters
for every interface property at runtime, so this.running = true
propagates through _lazyLoadDeviceState → getDeviceState proxy →
system state.
- Scripts plugin's mergeHandler auto-detects interfaces by mapping
method names: start/stop → StartStop, turnOn/turnOff → OnOff,
putSetting/getSettings → Settings, etc.
In theory the previous commit was sufficient. To pin down whether the
panel is calling something different (older Scrypted UIs lean toward
OnOff), this commit:
- Implements OnOff alongside StartStop. Both pairs are bound to the
same drain/bootstrap logic; whichever the panel calls, the user's
intent succeeds.
- Initialises this.running = false and this.on = false synchronously
in the constructor so the device state record carries a defined
value at registration time, instead of Scrypted defaulting to
Unknown on undefined.
- Logs 'lifecycle: X() called (...)' on every method entry. After
the user reloads + clicks STOP / START, the console will reveal
exactly which method Scrypted is invoking — or whether neither is.
If one of the methods fires, the other is dead code we can drop. If
neither fires on the panel's STOP/START button, the panel is a
Scripts-runtime control unrelated to device interfaces and we need
to look for a different lifecycle hook.
|
|
|
|
Silent early-return at 'if (!v.cameraId) return;' makes a brand-new
viewport with no camera selected look identical (in the console) to
one that subscribed successfully — there's no positive or negative
signal until you try to fire a camera event. After observing a fresh
viewport produce zero output on a doorbell press, switching the
early-return to a warning that says 'no camera assigned — open
Settings and pick a camera; subscription skipped' so the missing
configuration becomes self-evident.
|
|
|
|
The Scripts 'Status and Controls' panel previously rendered Unknown
because the Provider exposed no lifecycle interface. STOP / START in
that panel didn't do anything useful — Scrypted's only option was
to unload the script entirely, which left the user's re-paste flow
relying on the constructor-time globalThis cleanup hack (and even
that didn't cover in-flight streams until 1ce49d3).
This commit makes the panel real:
class ScryptedViewportProvider ... implements ..., StartStop
- async start() : drain shutdown cleaners, bootstrap, running=true
- async stop() : drain shutdown cleaners, clear viewport+listener+
stream maps, running=false
Both methods are idempotent (no-op when already in target state).
The constructor calls start() automatically so script load still
bootstraps without user action.
The private start() method that did the actual provisioning is
renamed to bootstrap() to free up the public name for the StartStop
contract. registerShutdownCleaners is renamed to drainShutdownCleaners
and now takes a 'reason' label so the log line distinguishes 'start:
tearing down N resources' (previous load) from 'stop: tearing down N
resources' (user-driven).
Effect on the workflow the user described:
1. Click STOP in the Scripts UI → stop() drains every resource
(streams, listeners, intervals) the provider currently holds.
2. Re-paste the new script.
3. New constructor runs → start() drains anything left behind →
bootstrap re-discovers children + re-attaches listeners → ready.
No more accumulating ghosts of prior loads.
|
|
|
|
Scrypted's Scripts sandbox doesn't release a previous load's resources
before constructing a new Provider instance, so anything long-lived
survives a re-paste. Previously we hand-rolled cleanup via two
separate globalThis hooks: __viewportListenerCleaners (camera event
listeners) and __viewportRegisterInterval (the 5-minute re-register
timer). Anything else — in-flight ffmpeg children, TCP sockets to
the firmware, the per-stream stats interval — leaked across reloads.
Visible symptom: re-pasting the script during an active stream left
the old stream's ffmpeg + socket running against a dead Provider
instance, accumulating across reloads until the host was restarted.
Unify all cleanup into one globalThis array, __viewportShutdownCleaners,
that every long-lived resource pushes onto. The next start() snapshots
and drains it before doing anything else. Resources covered:
- 5-minute re-register interval (clearInterval)
- Every camera + child-device event listener (reg.removeListener)
- Every in-flight stream's AbortController (abort.abort, which the
existing abort listeners chain into socket.destroy() + ffmpeg
SIGTERM + clearInterval on the stats logger)
Each cleanup is removed from the array when the resource naturally
ends (stream timeout / stopStream), so the list stays compact across
many stream cycles rather than growing forever.
Snapshot-before-walk avoids the iteration-skip bug where a stream
abort fires its abort listener, which splices the closure out of the
same array we're iterating. Snapshot the array, clear the original,
then walk the snapshot.
Doesn't address the 'Unknown' status panel in Scrypted Scripts UI —
that's a separate cosmetic issue (we don't implement a lifecycle
interface that reports running state).
|
|
|
|
The Settings 'Status (live)' section reaches the device via two
sequential GETs. A transient socket-level failure (httpd worker
pool briefly saturated by the active stream connection, mid-reboot
window, network jitter) used to leave the page showing
'device: offline / unreachable (fetch failed)' even though the
device was fine — a refresh would clear it.
One automatic 250 ms-backoff retry on either GET removes the
sporadic false positive. The 3 s per-request timeout stays the
same, so the worst-case page latency on a genuinely offline device
is ~6.5 s (3 + 0.25 + 3) instead of 3 s — acceptable for a
deliberate Settings open.
|
|
|
|
Unifi doorbell cameras expose the bell button as a child device of
the camera (separate nativeId, its own BinarySensor interface), not
as a property of the camera itself. The previous code did cam.listen()
on the parent camera only, so motion + person events arrived (those
fire on the camera itself) but bell-press events never reached
handleCameraEvent.
HomeKit kept working because Scrypted's HomeKit bridge auto-syncs all
child devices; our plugin was silently the only consumer missing the
event. Confirmed by the trace log in 3b0ab73 staying silent across a
bell press while motion events still printed.
Fix: in attachListener, walk systemManager.getDeviceIds() and pick
any device with providerId === cam.id as an additional listen target.
Subscribe on all of them with the same iface list — listen() no-ops
on unsupported ifaces, so we don't have to introspect each child.
Tracking changed from Map<nativeId, EventListenerRegister> to
Map<nativeId, EventListenerRegister[]> so detach/script-reload cleanup
removes every listener, not just the camera's.
|
|
|
|
painted<sent gap
Document the architecture pivot. The per-frame budget table now
reflects the long-lived TCP stream socket on port 81 (replaced the
HTTP /frame pattern) and the new recv/decode/paint timings:
recv 14-18ms on the wire (pure body), dec ~5ms, paint <35us, with
decode-task idle ~30ms per frame waiting on recv.
The critical change wasn't a TCP tuning knob — it was splitting
recv off its own FreeRTOS task with a 3-buffer PSRAM ring and a
1-deep latest-frame slot to the decode/paint task (d1c8d45). The
single-task loop blocked the socket for 6ms during decode+paint,
which against the IDF-default 5760-byte window forced a stop-go
cycle that capped painted at ~17fps. Raising the window to 65535
made it worse (g2g grew to 17s) because lwIP RX couldn't drain
the bursts and there's no way to skip-oldest on the kernel queue.
The task split eliminated the coupling: recv-task drains
continuously regardless of decode timing, the kernel buffer stays
near-empty, and painted now matches sent at the source rate.
Backlog entries for body shrink and OTA marked done.
|
|
handle_client previously ran recv → decode → paint serially on one
FreeRTOS task. The kernel TCP buffer filled during decode+paint
(~6ms), and against the IDF-default 5760-byte window the sender
naturally stop-go-rate-limited to ~consumption. Raising the window
to 65535 (previous experiment) regressed g2g from ~100ms to 17s
growing unbounded — the sender pumped 45+ segments per round into
a kernel buffer the app couldn't drain in time, and there was no
way to skip-oldest on the kernel queue.
This commit decouples recv from decode+paint:
recv-task: owns the socket. Reads header + body into one of three
preallocated PSRAM body buffers. On body complete, swaps
the just-filled buffer into a 1-deep pending slot and
picks a free buffer for the next recv. If the slot
already held a frame (decode is slow), drops oldest in
place — mirror of the Scrypted-side skip-oldest from
e5acf93.
decode-task: waits on a binary semaphore. On signal, claims pending,
then decodes + paints without holding any shared lock.
Frees its prior buffer implicitly by overwriting
s_decode_idx on the next claim.
3 PSRAM body buffers (~3MB of 28MB free) ensure the invariant
{recv_idx, pending_idx, decode_idx} are pairwise distinct without
ever blocking recv. jpeg_decoder.c grew an alloc_input_buffer helper
+ jpeg_decoder_decode now takes an explicit input pointer so the
stream and http_api snapshot paths don't share scratch.
New stats:
- recv_dropped_oldest: per-window count of pending-slot overwrites
- decode_idle_min/avg/max_us: time decode-task spent waiting on signal
Measurement at IDF-default 5760 window, Unifi medium substream:
before split: recv_avg=32ms recv_max~44ms fps=22-26 (recv blocked
during 6ms decode+paint; chunk_max capped at 5760)
after split: recv_avg=17ms recv_max=18-37ms fps=21-29 steady,
decode_idle_avg=27-40ms (decode mostly waiting),
drop_oldest=0, painted at source rate
The bottleneck moved from 'decode+paint serializes recv' to the
wire's own send rate. Bigger windows are now safe (recv-task drains
continuously, can't bury us), but won't add fps until source rate
goes up — that's a separate conversation.
|
|
regressed g2g
Step-2 experiment from the iterative tuning plan: bumped
CONFIG_LWIP_TCP_WND_DEFAULT 5760 → 65535 (and matching SND_BUF +
RECVMBOX 6 → 16) under the hypothesis that the 4×MSS window
throttled receive throughput.
Instrumentation showed the window WAS the per-call throttle
(recv_chunk_max was exactly 5760 before, grew to 33-60KB after).
But end-to-end got dramatically worse:
- g2g exploded from ~100ms steady to 17 SECONDS growing unbounded
(1964 → 6010 → 10632 → 14725 → 17889ms over consecutive windows).
- painted dropped from ~24fps to 15-20fps.
- recv_avg rose from 32ms steady to 50-90ms with frequent
230-500ms outliers (the retransmit / RTO signature).
- chunk_min fell to 61 bytes, chunk_avg dropped from 3415 to ~2200
— fragmentation as lwIP RX struggled with the bigger bursts.
Root cause: the single recv→decode→paint task serializes; the
kernel buffer fills against ANY window during the ~6ms
decode+paint, but with 65535 bytes available the sender pumps
~45 segments before stopping vs ~4 before. That overruns the
lwIP RX path (RECVMBOX=16 mailboxes, default pbuf pool) → drops
→ retransmits → multi-frame stalls. And Scrypted has no way to
skip-oldest on the kernel queue, so frames just pile up on the
wire / in the ESP32 kernel buffer → g2g grows linearly with time.
Lesson: window size isn't the right knob until the receiver can
drain at line rate independent of decode+paint. Next iteration:
split recv into its own FreeRTOS task with a 1-deep latest-frame
slot to the decode/paint task — only then revisit window.
|
|
handleCameraEvent now logs iface/typeof/data/details for every event
the camera fires on BinarySensor/MotionSensor/ObjectDetector. Helps
diagnose why a Unifi doorbell press isn't waking the viewport while
the same event reaches HomeKit. Temporary — remove once root cause is
confirmed.
|
|
call/chunk stats, SO_RCVBUF probe)
Per-frame samples aggregated over the existing 30-frame window:
- queued_at_body_start (FIONREAD just before body recv loop): how much
of the frame the kernel already absorbed during the previous
decode+paint. Close to jpeg_len → wire delivered the full frame
while we were busy (we're decode/paint-bound). Much smaller →
wire is throttled (window or buffer too small to absorb a frame
in our paint window).
- recv_calls: number of recv() syscalls the body read needed per
frame. High → small chunks → window-throttled sender.
- recv_chunk min/avg/max: bytes returned per recv() return in the
window. Avg = window body bytes / total syscalls.
- SO_RCVBUF: one-shot getsockopt at accept, logged and stashed in
stats. Confirms whether sdkconfig values reached the build —
TCP_WND_DEFAULT discrepancies are otherwise invisible.
All surfaced in the windowed log and in /state JSON alongside the
existing recv/dec/paint/idle stats. No behavior change yet.
|
|
Hold at most one pending frame; new ffmpeg frames arriving while the
socket is backpressured replace the held one (drop oldest) instead of
queueing. On 'drain', flush the held frame. The in-flight write stays
out of our hands. Steady-state Node-side buffer drops from up to 20MB
to ~1 frame, eliminating the cap-and-reconnect loop as the primary
shedding mechanism — it's now an emergency safety net for stuck
connections that never fire 'drain'.
Capture-time timestamp instead of write-time so a held frame keeps its
true age and g2g reflects any pending-slot hold latency.
New log fields: drop-oldest=N (frames shed from the pending slot this
window) and bp=N% (share of ffmpeg frames in the window that found the
link backpressured at decision time).
|
|
Streams the raw .bin to the inactive ota_0/ota_1 slot via esp_ota_*, flips otadata, replies 200, reboots after 500 ms. Single-shot guarded by atomic_flag (409 on concurrent). CONFIG_BOOTLOADER_APP_ROLLBACK_ENABLE armed: new images boot pending-verify and ota_arm_healthy_timer marks them valid after 30 s of healthy uptime; otherwise the bootloader reverts on next reset. /state gains ota_state.
|
|
|
|
Two reasons the Status (live) section was showing
"offline / unreachable (fetch failed)" intermittently:
1. /state and /config were fetched in parallel, eating both available
httpd sockets simultaneously (Phase 2 dropped max_open_sockets
from 4 to 2). Any other inbound HTTP at the same instant would
either queue or error.
2. The 1.5s timeout was tight given the firmware now juggles a live
stream socket on port 81 (with occasional cap-flush reconnects)
alongside the httpd workers on port 80.
Sequence the two fetches and bump the timeout to 3s. Total worst-case
6s if both are slow; that's fine for a Settings page, far from
"feels offline."
|
|
|
|
backlog
User's intent: "we want to see what's going on right now, not all the
shit we throttled and backlogged in the buffer." Dropping new arrivals
when the buffer is full doesn't help — the already-queued frames still
get processed in order, painting 5+ seconds of stale content before
catching up to live.
Better: when the buffer exceeds the configured cap, blow away the
whole queue by destroying the socket. The firmware's accept_task
picks up the next reconnect within ~500ms and the pipeline restarts
with the freshest ffmpeg frame as seq=1. Trade a sub-second visual
gap for clearing the entire backlog instantly.
Per-viewport setting "Max Scrypted-side buffer (MB)" under Display.
Default 20 MB (≈5s buffering ceiling at our measured ~3.5 MB/s
ffmpeg → firmware throughput). Range 1-200. Lower for tighter g2g
at cost of more frequent reconnect gaps under sustained backpressure;
higher to tolerate longer backlogs before flushing.
flushCount tracked per-stream and surfaced in the unified log line
alongside drops, plus node_buf shown vs the cap so the user can see
how close they're running to the threshold.
Demux loop now: if !socketReady → drop. Else if writableLength >
cap*MB → log + sock.destroy() + increment flushes. Else normal write.
|
|
|
|
Adds two values to the unified stream log line:
node_buf=NKB≈Mms
node_buf bytes is sock.writableLength — bytes that socket.write()
accepted from us but the kernel send buffer hasn't pulled yet, so
they sit in Node's internal queue. These have NOT been sent over the
wire; the firmware can't have them.
Dividing by the current send rate (bytes/sec computed from the same
window's bytesSent) gives the buffer depth in milliseconds: how many
seconds of already-extracted-from-ffmpeg frames are stalled at the
script. When g2g shows 7s steady state and node_buf shows e.g.
5000-6000KB ≈ 6500ms, the conclusion is obvious — the buffering
isn't in the firmware or on the wire, it's in Scrypted's Node
process.
No behavior change; pure diagnostic.
|
|
|
|
sent-vs-painted log
#1 — Camera listener leak
Same Scripts-sandbox lifecycle gap that bit us with setInterval (fixed
in 521de7e): cam.listen() returns an EventListenerRegister with a
removeListener() method. The listener is held on the camera plugin's
side, not the script's, so when the old Provider instance gets GC'd
on script reload its registrations stay live with a dead callback.
Field symptom: after re-pasting the script N times, every camera
event arrives N times in handleCameraEvent. Today the streamStarting
guard catches sequential dupes, but two callbacks firing on the same
event within the JS microtask window both check streamStarting
BEFORE either has added their nativeId, both proceed, two
pushSnapshots launch in parallel. They race for the firmware
decoder lock — one paints, one gets 503 — and the user perceives
sharp-fighting-with-itself as "snapshot quality is bad again."
Fix: same globalThis trick as the setInterval. Push a remover
closure onto G.__viewportListenerCleaners every time attachListener
fires; at start() of the new Provider instance, drain the array
calling each remover before creating fresh listeners. Idempotent
across reloads.
#2 — Unified sent-vs-painted log
The user asked: "do we have an understanding in the logging/
instrumentation what the FPS of the stream output is from scrypted
out is vs the actual rendered FPS once the stream is loaded?"
Both numbers existed but in separate log lines (skipLogger every 10s
+ fwPoller every 5s, interleaved in the console). Folded the /state
poll into the 10s skipLogger so there's ONE line per window with
sent fps + painted fps side-by-side plus the gap explicitly labeled:
stream "kitchen": sent=24.2fps painted=22.8fps (fw-skipped=1.4fps,
drops=0) 4.58MB/s sent / 4.29MB/s painted |
socket.write p50=0ms p95=1ms max=5ms backpressured=true |
recv=27776/37237/44470us dec=5788/5991/6626us
paint=30/36/46us idle=164/588/11383us | g2g=142ms
fw-skipped = (sent − painted), how many frames per second the
firmware's FIONREAD skip dropped to keep the panel on the freshest
frame. The g2g value is now meaningful per-frame thanks to the
firmware-side live-update we landed last commit.
Drops the separate fwPoller; one comprehensive log per 10s window.
|
|
frame age
Previous Phase 4 implementation only published last_paint_event_us_low
into the windowed-stats snapshot at every 30-frame roll (~1.5s at
20fps). Combined with the script's 5s /state poll cadence, the
"freshest frame age" we could compute was up to ~6.5s stale before
any actual lag — meaningless as a perf signal.
Update s_stats.last_paint_event_us_low under portMUX on every
painted v1 frame. The other window stats (recv/dec/paint/idle
min/avg/max + fps/MBps) keep their roll cadence because they
genuinely need a full window to compute; this single u32 just gets
overwritten with the most recent value each paint.
portENTER_CRITICAL on the ESP32-P4 is ~half a microsecond per side
— at 20fps that's 20µs/s of overhead, immeasurable next to the
30-40ms recv per frame.
Expected: g2g during a saturated stream drops from the previous
2-6s reading to single-digit hundreds of ms or lower, reflecting
the actual emit→display lag.
|
|
|
|
Two follow-up fixes informed by the first-round field data:
#1 — g2g measures display age, not time-since-wake
Previously eventUsLow was stamped once at wake event and reused for
the whole stream session. /state's last_paint_event_us_low echoed
back the wake-time anchor, so g2g = "elapsed time since user wake."
That metric is true but not actually informative — it just grows
linearly with stream duration.
What's actually useful for the "is the panel showing current
reality" question is the AGE of the currently-displayed frame.
Stamp event_us_low per emitted frame at Date.now()*1000 (low 32
bits) inside the stream demux loop. Now last_paint_event_us_low
reflects the most recent painted frame's emit time, and
g2g = (script_now_us_low - last_paint_event_us_low) = display
staleness in ms.
#2 — sharp first-paint dropped from ~1800ms to ~800ms
Previous commit (phase 5) turned on sharp's mozjpeg encoder for
slightly tighter file size at the same JPEG quality. Field
measurement showed it costs ~1.5s of CPU per snapshot vs
libjpeg-turbo's ~400ms — a ~4× regression on the metric we care
most about (event → first-paint). The file-size win is ~5%; on
800x480 panel output that's not perceptible.
Drop mozjpeg: true. Quality settings (quality=100, chroma 4:4:4 at
jpegQuality ≤ 2) stay the same, so visual output is unchanged.
First-paint comes back down to where it should be.
|
|
stream_server
Two pieces of documentation, neither changes behavior:
1. TESTING.md gains a "Performance review playbook" section: per-
session capture commands for firmware serial + /state poll +
Scrypted console, an annotated walkthrough of what each log
shape means, an investigation-threshold table mapping
user-facing symptoms to likely causes and the first thing to
check, and a one-line tools list. Replaces the implicit "ask the
maintainer how to debug" loop with a reproducible workflow.
2. stream_server.c gains a "Why TCP and not UDP" header comment
documenting the analysis from the design phase. JPEGs are ~123
IP datagrams at 1500 MTU; on hardwired Gigabit LAN switch-fabric
loss is < 1e-9/packet → per-frame corruption ≈ 1.2e-7. UDP's
theoretical wins (no Nagle, latest-wins semantics) don't apply
because TCP_NODELAY is on, socket.write p50 < 1ms, and the
FIONREAD trick already implements latest-wins on the receive
side. UDP's costs (200-400 LOC of app-layer fragmentation, loss
of nc/curl debug, FIONREAD trick stops working under
fragmentation) are real. Documented as reference so a future
contributor doesn't re-derive the analysis from scratch.
|
|
|
|
Wire format change: stream frames now carry a 4-byte "VPRT" magic +
4-byte jpeg_len + 4-byte seq + 4-byte event_us_low. Total 16 bytes
(was 8). The firmware sniffs the first 4 bytes per frame: if they
spell VPRT it reads the remaining 12 bytes of v1 header; otherwise
it interprets bytes 0-3 as jpeg_len for the old v0 8-byte format and
reads 4 more for seq. Lets a v1 firmware accept a v0 (legacy)
Scrypted script during the rollout window. v0 will be removed once
all field deployments roll forward.
event_us_low is the low 32 bits of the Scrypted host's monotonic µs
at camera-event arrival. The firmware does NOT interpret it (the
clocks aren't sync'd); it just stamps it on every painted frame and
exposes the most recent value via /state. The script polls /state
every 5s during an active stream, reads last_paint_event_us_low,
and computes glass-to-glass = (now_us_low - last_paint_event_us_low)
with 32-bit wrap. 30s sanity ceiling on the wrap to discard event
timestamps from before the stream started.
Also expose the firmware's just-closed 30-frame window stats via
/state under the "stream" key — frames, bytes, window_us, plus
min/avg/max for recv/dec/paint/idle. Lets external tools (a curl
loop, the Scrypted plugin, etc) poll the firmware's view without
parsing serial logs.
Firmware:
- stream_server.h: 16-byte v1 wire spec, stream_server_stats_t
struct, stream_server_snapshot_stats(out) getter.
- stream_server.c: magic-detect header read path, last_event_us_low
capture into per-connection state, portMUX-protected window-stats
snapshot at every 30-frame roll.
- http_api.c: GET /state JSON gains a "stream" sub-object with the
full snapshot.
- viewport_state.h: VIEWPORT_VERSION 1.0.0 → 1.1.0 (new /state shape).
Scrypted:
- startStream captures eventUsLow = (tEvent * 1000) >>> 0.
- TCP demux loop writes the 16-byte v1 header with the VPRT magic.
- New fwPoller setInterval (5s) fetches /state, parses .stream,
computes g2g, emits one summary line per poll cycle.
|
|
|
|
User reported "snapshot is still lower qual than the stream" even at
jpegQuality=1. Two compounding causes in the sharp transform path:
1. Quality mapping topped out at 97. Previous formula was
Math.max(50, 100 - v.jpegQuality * 3); at jpegQuality=1 that
produced quality=97 — versus the stream's ffmpeg mjpeg -q:v 1
which lands at sharp-equivalent ~99-100. Visible compression
artifacts on flat regions accounted for ~half the perceived
quality gap. New formula Math.min(100, 102 - v.jpegQuality * 2)
yields:
jpegQuality=1 → 100 (was 97)
jpegQuality=2 → 98 (was 94)
jpegQuality=5 → 92 (was 85)
jpegQuality=10 → 82 (was 70)
jpegQuality=31 → 40 (was 50, clamped)
2. Chroma subsampling was sharp's default 4:2:0 — half-rate chroma
in U+V planes. On colored edges (UI overlays, sharp camera
subjects against backgrounds, text) this smears and is the
dominant visible artifact at panel-native 800x480. ffmpeg
mjpeg's default behavior at -q:v 1 is closer to 4:2:2 or no
subsampling. Force 4:4:4 when jpegQuality <= 2 (the "I want max
quality" regime); keep 4:2:0 for jpegQuality >= 3 since the
point at that quality is smaller files.
Plus mozjpeg: true on the encoder. Same quality value, slightly
tighter file size — modest but free.
The Lanczos kernel was already lanczos3 — confirmed, no change.
Expected outcome: snapshot at jpegQuality=1 is visually
indistinguishable from a stream frame at the same scene.
|
|
|
|
Adds per-stage timing logs to both the stream and snapshot paths,
all relative to a single t=0 anchor: the moment the camera event
arrived in handleCameraEvent. Lets us see whether snapshot and
stream actually start in parallel (they should, within ~2ms) and
where wall-clock goes inside each path.
handleCameraEvent now captures tEvent = Date.now() and threads it
into startStream(v, tEvent). startStream threads it further into
pushSnapshot(v, cam, tEvent). Both helpers define a `since()` arrow
that computes the current offset.
New log lines on wake:
event MotionSensor -> "kitchen": fired at +0ms (wake)
stream "kitchen": start +0ms
stream "kitchen": socket connect requested +1ms
snapshot "kitchen": start +1ms
snapshot "kitchen": takePicture +320ms
stream "kitchen": socket connect open +24ms
snapshot "kitchen": transform +458ms via sharp (217KB)
snapshot "kitchen": post sent +459ms
snapshot "kitchen": post acked +551ms ← first user-visible paint
stream "kitchen": first ffmpeg frame +780ms (jpeg=210KB)
stream "kitchen": first socket.write +781ms
post_acked is the snapshot's true glass-to-glass: /frame returns
after display_flip_back_buffer, so by the time the response
resolves the firmware has the new pixels queued for the next DPI
scanout. No separate firmware-side wiring needed for the snapshot
path's g2g measurement.
The stream path's true g2g still needs the 16-byte header
extension in Phase 4 — `first socket.write` is just "Scrypted sent
the bytes," not "panel showed them."
|
|
|
|
Both sides simultaneously because the script's X-Frame-Seq sender and
the firmware's X-Frame-Seq receiver negotiated a contract that's now
retired entirely.
Firmware (main/http_api.c):
- Delete static uint32_t s_last_painted_seq (declaration + reset in
state_post_handler + the read in frame_post_handler + the
assignment after paint).
- Delete the X-Frame-Seq header read (httpd_req_get_hdr_value_str
+ strtoul block).
- Delete the X-Frame-Drop: stale-seq response path.
- Delete the entire Server-Timing httpd_resp_set_hdr block.
- Delete the setsockopt(TCP_NODELAY) at frame_post_handler entry.
/frame is a single body POST → empty 204; no second packet for
Nagle to coalesce with on the response side.
- Delete static int64_t s_last_post_us + the idle_us computation.
/frame fires at most once per wake; idle gap between wakes is
dominated by user/event timing, not anything firmware-controllable.
- Drop the per-10-frames condition on the timing log — /frame fires
rarely enough that one log per snapshot is the right cadence —
and rename "frame N: ..." → "snapshot: ..." to match.
- Drop cfg.max_open_sockets 4 → 2 (snapshot POST + concurrent /state
or /config).
- Drop #include "lwip/sockets.h" (orphaned with TCP_NODELAY).
Script (scrypted/scrypted-viewport.ts):
- Delete the agents Map + agentFor() method (per-host keepAlive
Agent pool; over-engineered for ~1 POST/min control plane).
- Delete httpRequest() helper (40 lines wrapping http.request to
surface tHeaders/tDone — no consumer reads those fields anymore).
- Rewrite postJSON() to a 10-line fetch() with AbortSignal.timeout.
- Delete the frameSeq Map + the seq counter + X-Frame-Seq header on
the snapshot fetch + the X-Frame-Drop response check + the
frameSeq.delete in stopStream.
- Extract buildVf(orientation, panelW, panelH) helper near top of
ScryptedViewportProvider — used by both startStream (live) and
pushSnapshot (one-shot ffmpeg fallback).
TCP_NODELAY references remain in the live-stream socket path
(scrypted-viewport.ts:773, 788). That's a different socket (raw
net.Socket on TCP/81) and noDelay there is what eliminates Nagle
stalls under the streaming-heavy live workload. Load-bearing.
|
|
|
|
Removes/rewrites every comment that referenced the dead HTTP-streaming
architecture so a reader of v1.0.0+ source isn't chasing a model
that's been gone for several releases. No code paths changed; this is
the safe pre-pass before Phase 2's actual deletions.
Firmware (main/http_api.c):
- s_last_painted_seq / s_last_post_us / X-Frame-Seq parse / stale-
frame guard / Server-Timing emission / TCP_NODELAY setsockopt /
max_open_sockets=4 — each block now leads with "(Legacy from the
HTTP-streaming era; removed in Phase 2)" so the reader knows the
block is doomed, not load-bearing.
- /frame dim-mismatch comment narrowed: panel-native is always 800x480
BGR888, no Scrypted-side variation expected; Scrypted does the
rotation+scale via sharp/mediaManager/ffmpeg cascade (snapshot) or
ffmpeg -vf (stream).
- "Single in-flight frame. Concurrent posts get 503" rewritten to
reflect: decoder mutex now mostly serves to fence /frame snapshots
against an active stream_server decode.
Firmware (main/stream_server.c:136-142):
- FIONREAD-skip rationale reduced from a 7-line paragraph to one
sentence; the savings/tradeoff math now lives in the plan, not the
per-line comment.
Script (scrypted/scrypted-viewport.ts):
- Top-of-file tuning constants block drops the "frame_interval_ms
removed", "fps filter", "in-flight back-to-back startStream"
rationale; one short line covers the model: "Stream rate is paced
by camera + TCP backpressure; no app-level fps cap."
- agentFor() rationale rewritten: this Agent is over-engineered for
control-plane traffic (~1 POST/min steady state) — a legacy of when
it backed per-frame /frame POSTs. Marked for Phase 2 retirement.
- noDelay on Agent: clarified it's now a no-op safety for control
plane (was load-bearing for live-stream pipelining).
- snapshot fire-and-forget comment: replaced X-Frame-Seq-race
rationale with the actual TCP-streaming truth (sharp/mediaManager/
ffmpeg cascade race against stream socket bring-up).
- writeLatencies probe: rewrote the "keep-alive socket" comment
(live stream uses raw net.Socket, not the http.Agent pool).
- socketBackpressured: explicit "diagnostic only, never gates writes"
comment added at declaration site.
- skipLogger header: rewrote the inFlight/MAX_INFLIGHT/fps-filter
rationale into a one-line description of what the log line actually
emits today.
- frameSeq map: now flagged as legacy of HTTP-streaming era, retired
in Phase 2.
|
|
First milestone where the device runs the way it was intended end-
to-end:
- Raw-TCP streaming data plane at ~14-20 fps painted, sub-ms socket
writes, no Nagle/ACK pathology.
- HTTP control plane (/state, /config) plus HTTP /frame for the
one-shot snapshot fast path on wake. Snapshot via sharp → ~700ms
first paint from a cold trigger (down from ~1000ms+ via ffmpeg).
- Firmware-side "always paint the latest" — FIONREAD skip drops
superseded frames before decode, keeping glass-to-glass tight
even when upstream produces faster than we can ingest.
- HW JPEG decode + zero-copy paint into the panel back framebuffer.
- 12 fps goal hit cleanly at q:v 1 (visually lossless). Source
pinned to medium-resolution substream so the camera supplies
enough frames to actually fill the pipe.
- Scrypted side: per-host node:http keep-alive Agent with NODELAY,
the cascading sharp → mediaManager → ffmpeg snapshot transform,
wake-trigger picker on the new-device dialog, and the leaked-
setInterval-across-script-reloads bug squashed.
Bumping VIEWPORT_VERSION 0.1.0 → 1.0.0 so the boot log and info
screen reflect it.
|
|
This reverts commit 6e36cbc14a06b8c0e15bfe78afb8adb43f8909ba.
|
|
Hypothesis confirmed by the symptoms: with WND_DEFAULT=65535,
SND_BUF=65535, RECVMBOX=16, the firmware boots fine and the stream
server logs "listening on tcp/:81", but Scrypted's first SYN never
produces a "client connected from ..." entry. The script-side
socket stays stuck in pending-connect — neither a connect nor an
error event fires — and every ffmpeg-emitted frame gets dropped.
Root cause is almost certainly lwIP's PBUF / MEMP pools not being
scaled to back the 64KB windows on multiple sockets (httpd has 4,
stream has 1, that's 5 × 64KB = 320KB of buffer commitment from a
heap that's also feeding PSRAM-backed framebuffers, JPEG scratch,
and Ethernet DMA). lwIP doesn't surface the allocation failure as
an accept error — it just refuses to complete the handshake.
Backing off the changes for now. The streaming path already hits
~14-20 fps at our measured 5.3 MB/s ingest ceiling under the lwIP
defaults — that's plenty above the 12-fps goal. Revisit window
tuning later by also bumping CONFIG_LWIP_PBUF_POOL_SIZE and
MEMP_NUM_TCP_PCB so the buffer commitment is actually allocatable.
Keeping the FIONREAD skip in stream_server (it's orthogonal to lwIP
buffer sizes and only helps).
|
|
|
|
backpressure-blind + lwIP TCP window bump
Three changes that work together to make the stream "always paint
what's freshest, never sit on a stale frame":
#1 — Firmware FIONREAD skip in stream_server
Right after the body of frame N comes off the wire (and before we
unlock the decoder + spend ~6ms on decode + paint), check the
kernel receive buffer with ioctl(FIONREAD). If at least one more
header (8 bytes) is queued, frame N is no longer the freshest
possible — skip its decode + paint and loop back to read frame
N+1. The TCP recv cost is unavoidable (bytes still have to cross
the wire) but the decoder + paint cost is saved on every superseded
frame. Glass-to-glass latency on the latest frame drops by however
many frames had backed up.
#2 — Scrypted: keep writing past kernel-buffer backpressure
Previously: when sock.write() returned false we dropped the next
ffmpeg frame at source. New: we keep writing through. Node buffers
internally; under our load (~16 MB/s ffmpeg → ~5-7 MB/s firmware)
the buffer rarely exceeds a frame or two. With the firmware now
silently skipping decode on backed-up frames (#1), excess frames
get shed for free on the device side. Scrypted's job is just to
hand the firmware the freshest bytes as fast as possible.
socketBackpressured is still tracked for the diagnostic log.
#3 — lwIP TCP window bump (revisiting earlier regression)
The previous attempt at LWIP_TCP_WND_DEFAULT=32k regressed under
HTTP because every /frame opened a fresh socket and we paid the
slow-start cost repeatedly. The streaming pivot eliminated that:
the socket is long-lived, slow-start runs exactly once, then we
ride the full window for the rest of the session. Bumping to
65535 (max for stock lwIP), SND_BUF to match, RECVMBOX to 16, SACK
on. Expected: recv throughput ceiling moves up from ~5.3 MB/s,
which directly raises the fps ceiling (recv is currently 37ms of
the 43ms per-frame total).
|
|
|