Files
otelsetup/otel_test.go
T
argoyleandClaude Opus 5 f1551a5f47
otelsetup / test (push) Skipped
otelsetup / vulnerabilities (push) Skipped
pre-commit / pre-commit (push) Skipped
otelsetup / vulnerabilities (pull_request) Successful in 1m0s
otelsetup / test (pull_request) Successful in 1m21s
pre-commit / pre-commit (pull_request) Successful in 3m36s
fix: close idle OTLP connections before the collector does
Metrics are pushed every 60s and Alloy's OTLP receiver closes connections idle for 1m, while the exporters keep them for 90s. A push could reuse a connection the collector was closing and fail with EOF or connection reset; the data was dropped (no retry for a POST or for transport errors). In prod that was ~3 failed metric pushes per pod per hour, plus occasional trace batches.

Both exporters now use an HTTP client that closes idle connections after 30s. A custom client makes the exporters ignore the OTLP timeout and certificate env vars; none are set.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014rfQ5HJ7zwfuYQcWXd3Mm3
2026-09-19 15:43:14 +02:00

23 lines
625 B
Go

package otelsetup
import (
"net/http"
"testing"
"time"
)
func TestOTLPHTTPClient_ClosesIdleConnectionsBeforeTheCollector(t *testing.T) {
c := otlpHTTPClient()
tr, ok := c.Transport.(*http.Transport)
if !ok {
t.Fatalf("transport is %T, want *http.Transport", c.Transport)
}
// Alloy's OTLP receiver closes connections idle for 1m; the client must close first.
if tr.IdleConnTimeout <= 0 || tr.IdleConnTimeout >= time.Minute {
t.Errorf("IdleConnTimeout = %v, want in (0, 1m)", tr.IdleConnTimeout)
}
if c.Timeout != 10*time.Second {
t.Errorf("Timeout = %v, want 10s (the exporters' default)", c.Timeout)
}
}