fix: close idle OTLP connections before the collector does #184

Merged
argoyle merged 1 commits from fix-otlp-idle-connections into main 2026-09-19 13:51:38 +00:00
Owner

Fixes the prod failed to upload metrics: Post ".../v1/metrics": EOF / connection reset by peer / http: server closed idle connection errors.

Cause. Services push metrics every 60s (PeriodicReader default). Alloy's otelcol.receiver.otlp closes HTTP connections that have been idle for 1m (idle_timeout default). The exporters' own transport keeps idle connections for 90s, so a push could reuse a connection the collector was already closing. net/http doesn't retry a POST, and the exporters don't retry transport errors, so that push was dropped.

  • Over 6h in prod: 595 failed metric pushes across 14 services and 33 pods (~3 per pod per hour), plus 26 failed trace batches.
  • Alloy itself is healthy (no restarts, low CPU) and the failures aren't clustered around its events.
  • A scaled local repro (server idle 100ms, pushes every 100ms) failed 138 of 300 pushes with client idle > server idle, and 0 of 300 with client idle < server idle.

Fix. Both exporters get WithHTTPClient(otlpHTTPClient()). That is a copy of the exporters' default transport (v1.46.0 ourTransport) with IdleConnTimeout 30s, below the collector's 1m. The client timeout is 10s, the exporters' default.

Trade-off: a custom client makes the exporters ignore OTEL_EXPORTER_OTLP_*TIMEOUT and the TLS certificate env vars. No service sets them (all only set OTEL_EXPORTER_OTLP_ENDPOINT). This is noted in the code comment and CLAUDE.md.

Impact of the bug: temporality is cumulative, so counters and histograms stayed correct. It cost a missing sample about 5% of the time and dropped some trace batches.

Review (Go Backend Expert): no Critical or High findings. It confirmed the transport matches upstream apart from the idle timeout, and that headers, compression and retry are unaffected. I applied its Mediums: the same fix for the trace exporter (confirmed failing in prod), and documenting the ignored env vars. Accepted gap: the unit test pins the invariant (client idle < 1m, timeout 10s) but not the WithHTTPClient wiring; an end-to-end test would need a fake clock. I skipped closing idle connections on shutdown, since the 30s timeout reaps them anyway.

Tests: go test -race ./... passes; prek is clean.

Rollout: after the release PR, bump all 14 services from v0.6.0.

🤖 Generated with Claude Code

https://claude.ai/code/session_014rfQ5HJ7zwfuYQcWXd3Mm3

Fixes the prod `failed to upload metrics: Post ".../v1/metrics": EOF` / `connection reset by peer` / `http: server closed idle connection` errors. **Cause.** Services push metrics every 60s (PeriodicReader default). Alloy's `otelcol.receiver.otlp` closes HTTP connections that have been idle for 1m (`idle_timeout` default). The exporters' own transport keeps idle connections for 90s, so a push could reuse a connection the collector was already closing. net/http doesn't retry a POST, and the exporters don't retry transport errors, so that push was dropped. - Over 6h in prod: 595 failed metric pushes across 14 services and 33 pods (~3 per pod per hour), plus 26 failed trace batches. - Alloy itself is healthy (no restarts, low CPU) and the failures aren't clustered around its events. - A scaled local repro (server idle 100ms, pushes every 100ms) failed 138 of 300 pushes with client idle > server idle, and 0 of 300 with client idle < server idle. **Fix.** Both exporters get `WithHTTPClient(otlpHTTPClient())`. That is a copy of the exporters' default transport (v1.46.0 `ourTransport`) with `IdleConnTimeout` 30s, below the collector's 1m. The client timeout is 10s, the exporters' default. **Trade-off:** a custom client makes the exporters ignore `OTEL_EXPORTER_OTLP_*TIMEOUT` and the TLS certificate env vars. No service sets them (all only set `OTEL_EXPORTER_OTLP_ENDPOINT`). This is noted in the code comment and CLAUDE.md. **Impact of the bug:** temporality is cumulative, so counters and histograms stayed correct. It cost a missing sample about 5% of the time and dropped some trace batches. **Review (Go Backend Expert):** no Critical or High findings. It confirmed the transport matches upstream apart from the idle timeout, and that headers, compression and retry are unaffected. I applied its Mediums: the same fix for the trace exporter (confirmed failing in prod), and documenting the ignored env vars. Accepted gap: the unit test pins the invariant (client idle < 1m, timeout 10s) but not the `WithHTTPClient` wiring; an end-to-end test would need a fake clock. I skipped closing idle connections on shutdown, since the 30s timeout reaps them anyway. Tests: `go test -race ./...` passes; prek is clean. Rollout: after the release PR, bump all 14 services from v0.6.0. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_014rfQ5HJ7zwfuYQcWXd3Mm3
argoyle added 1 commit 2026-09-19 13:44:48 +00:00
fix: close idle OTLP connections before the collector does
otelsetup / test (push) Skipped
otelsetup / vulnerabilities (push) Skipped
pre-commit / pre-commit (push) Skipped
otelsetup / vulnerabilities (pull_request) Successful in 1m0s
otelsetup / test (pull_request) Successful in 1m21s
pre-commit / pre-commit (pull_request) Successful in 3m36s
f1551a5f47
Metrics are pushed every 60s and Alloy's OTLP receiver closes connections idle for 1m, while the exporters keep them for 90s. A push could reuse a connection the collector was closing and fail with EOF or connection reset; the data was dropped (no retry for a POST or for transport errors). In prod that was ~3 failed metric pushes per pod per hour, plus occasional trace batches.

Both exporters now use an HTTP client that closes idle connections after 30s. A custom client makes the exporters ignore the OTLP timeout and certificate env vars; none are set.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014rfQ5HJ7zwfuYQcWXd3Mm3
argoyle scheduled this pull request to auto merge when all checks succeed 2026-09-19 13:44:50 +00:00

Coverage Report

Total coverage: 23%

## Coverage Report Total coverage: **23%**
argoyle merged commit 8b3d3110fe into main 2026-09-19 13:51:38 +00:00
argoyle deleted branch fix-otlp-idle-connections 2026-09-19 13:51:39 +00:00
Sign in to join this conversation.
No Reviewers
2 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: shiny/otelsetup#184