otelsetup / test (push) Skipped
otelsetup / vulnerabilities (push) Skipped
pre-commit / pre-commit (push) Skipped
otelsetup / vulnerabilities (pull_request) Successful in 1m0s
otelsetup / test (pull_request) Successful in 1m21s
pre-commit / pre-commit (pull_request) Successful in 3m36s
Metrics are pushed every 60s and Alloy's OTLP receiver closes connections idle for 1m, while the exporters keep them for 90s. A push could reuse a connection the collector was closing and fail with EOF or connection reset; the data was dropped (no retry for a POST or for transport errors). In prod that was ~3 failed metric pushes per pod per hour, plus occasional trace batches. Both exporters now use an HTTP client that closes idle connections after 30s. A custom client makes the exporters ignore the OTLP timeout and certificate env vars; none are set. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014rfQ5HJ7zwfuYQcWXd3Mm3
23 lines
625 B
Go
23 lines
625 B
Go
package otelsetup
|
|
|
|
import (
|
|
"net/http"
|
|
"testing"
|
|
"time"
|
|
)
|
|
|
|
func TestOTLPHTTPClient_ClosesIdleConnectionsBeforeTheCollector(t *testing.T) {
|
|
c := otlpHTTPClient()
|
|
tr, ok := c.Transport.(*http.Transport)
|
|
if !ok {
|
|
t.Fatalf("transport is %T, want *http.Transport", c.Transport)
|
|
}
|
|
// Alloy's OTLP receiver closes connections idle for 1m; the client must close first.
|
|
if tr.IdleConnTimeout <= 0 || tr.IdleConnTimeout >= time.Minute {
|
|
t.Errorf("IdleConnTimeout = %v, want in (0, 1m)", tr.IdleConnTimeout)
|
|
}
|
|
if c.Timeout != 10*time.Second {
|
|
t.Errorf("Timeout = %v, want 10s (the exporters' default)", c.Timeout)
|
|
}
|
|
}
|