Documentation
¶
Overview ¶
Package server exposes the metadata store over HTTP.
Index ¶
Constants ¶
This section is empty.
Variables ¶
var ErrSecretKeyNotFound = errors.New("key not in secret")
ErrSecretKeyNotFound reports that the Secret exists but does not carry the requested key. SecretReader implementations must wrap it (%w) for that case, so handleScrapeAuth can answer 404 — a client-caused miss — and reserve the retryable 502 for failures that are the CLUSTER's: a forbidden read, a timeout, an unreachable API server. Collapsing the two made an RBAC denial indistinguishable from a typo in a monitor's secret ref.
Functions ¶
func Enrich ¶
func Enrich(r MetadataResolver, pod *kubemeta.Pod, refs []metav1.OwnerReference)
Enrich fills in owner-chain and namespace metadata on a pod document: the ONE enrichment every served pod gets, used by the HTTP handlers and by the service's own self-attribute lookup (cmd/kubescrape selfResolver), so an enrichment step added here reaches both. Package-level rather than a method because that caller holds a MetadataResolver, not a Server.
Types ¶
type Config ¶
type Config struct {
Store *store.Store
Services *services.Index
// Monitors serves ServiceMonitor-derived targets (nil = disabled).
Monitors *servicemonitors.Index
Resolver MetadataResolver
// OwnerGeneration is the change token of the owner and namespace informer
// caches Resolver reads (owners.Changes.Generation). It cannot come from
// Resolver itself: a Resolver holds no state to notice a change in, and the
// informers that do are registered before it exists.
//
// nil DISABLES the node-targets ETag memo's token path, falling back to the
// wall clock — deliberately, because an unwired source is indistinguishable
// from one that never changes, and trusting its constant zero would serve a
// frozen target list. Wire it or accept the slower path; do not let it
// default to something that looks valid.
OwnerGeneration func() uint64
// MaxWait is the default and maximum time a container lookup may block
// waiting for metadata to appear. Requests may shorten it with ?wait=.
MaxWait time.Duration
// CacheTTL sets the max-age on metadata responses (Cache-Control + ETag),
// letting the agent's HTTP client serve repeat lookups locally. 0 disables
// cache headers.
CacheTTL time.Duration
// Ready is closed once the informer caches have synced.
Ready <-chan struct{}
// Secrets serves monitor endpoints' bearer-token Secrets to agents
// (GET /v1/scrape-auth/...); nil disables the endpoint (404). Opt-in via
// -scrape-auth-secrets — it requires secrets RBAC and ships secret
// material over the cluster-internal HTTP channel.
Secrets SecretReader
// ScrapeAuthTokens yields the bearer tokens clients may present on
// GET /v1/scrape-auth (`Authorization: Bearer <token>`); cmd/kubescrape
// wires it from -scrape-auth-token-file through bearer.Rotating. It is
// REQUIRED whenever Secrets is set (see Validate) and guards that route
// only — the rest of the API carries no secret material and stays open.
// It is evaluated per request and every returned token is accepted, which
// is what makes ROTATION a non-event: the source returns the current token
// plus the previous one for a grace window, so re-reading agents and the
// re-read service file never have to flip in lockstep. An empty token never
// authorizes (bearer.Authorized).
ScrapeAuthTokens func() []string
// Log receives the handful of server-side events an agent cannot diagnose
// from a status code alone (a Secret read that failed for a reason other
// than "no such key"). nil uses slog.Default().
Log *slog.Logger
}
Config configures the HTTP server.
func (Config) Validate ¶
Validate reports a configuration that must not start the process.
The one rule: a Server that can read Secrets must know the token guarding them. Without it GET /v1/scrape-auth would serve every Secret key a monitor endpoint references to anything that can reach the service — a silent, cluster-wide secret leak, which is exactly the failure mode a "the flag was not set" default must never produce.
type MetadataResolver ¶
type MetadataResolver interface {
Resolve(namespace string, refs []metav1.OwnerReference) ([]kubemeta.Owner, int)
Namespace(name string) *kubemeta.ObjectMeta
Node(name string) *kubemeta.ObjectMeta
}
MetadataResolver enriches pods with related-object metadata: the full owner chain, the pod's namespace metadata and node metadata.
Resolve's second result is how many ownerReferences its per-pod ceiling refused to describe (owners.MaxOwners); it rides the served document as kubemeta.Pod.OwnersOmitted so a truncated chain cannot read as a complete one.
type SecretReader ¶
type SecretReader interface {
Get(ctx context.Context, namespace, name, key string) (string, error)
}
SecretReader resolves one Secret key's value.
type Server ¶
type Server struct {
// contains filtered or unexported fields
}
Server serves container metadata and node scrape targets.
func (*Server) Drain ¶
Drain releases every request parked waiting for metadata and refuses the later ones, returning how many were released. Idempotent; safe to call from any goroutine.
It covers BOTH parking spots a container lookup has, which is the whole point: the store's per-ID waiter (Store.Drain) and this Server's wait for the initial informer sync (waitReady). The second one is easy to forget and is the one that bites on the worst path — a SIGTERM arriving before the caches sync leaves no store waiters at all, because no request has got that far, so draining only the store would still leave every lookup parked for its full wait budget and cut without a response when the process exits.
It must run BEFORE http.Server.Shutdown: Shutdown WAITS for these handlers and, at its deadline, returns without closing anything, so a request left parked is cut by the process exit rather than answered.
Idempotent means the COUNT too, not just the close: a second call reports 0. The whole body is inside the guard because the number is a log line's — the caller logs "released blocked container lookups" when it is nonzero — and a readyParked read left outside would let a second Drain re-report parks the first call had already released and re-log the loss they represent. The store's half returns 0 on its own second call; this keeps the two halves answering the same way.
func (*Server) HTTPServer ¶
HTTPServer wraps s.Handler() in an http.Server with hardened timeouts. The metadata service fronts a whole DaemonSet fleet, so a single slow, buggy or hostile client must never pin connections and goroutines indefinitely:
- ReadHeaderTimeout kills Slowloris-style header trickling.
- IdleTimeout reaps parked keep-alive connections.
- ReadTimeout and WriteTimeout bound trickled request bodies (which the handlers never read, but net/http drains before connection reuse) and stuck response writes. Both MUST exceed MaxWait: the container endpoint legitimately holds a request for up to MaxWait, WriteTimeout's clock starts when the request headers are read, and a ReadTimeout shorter than the handler's runtime cancels the request context (net/http's background read hits the whole-request read deadline mid-handler), which would abort legitimate waits early. Hence MaxWait + slack, never a fixed constant.
- MaxHeaderBytes bounds the wire size of a head; releaseParkedHead is what makes a parked request's COST bounded rather than only its count.