Logging operator
Configure and operate the kube-logging logging-operator (formerly Banzai Cloud) on Kubernetes — the CRD-driven log pipeline: Fluent Bit collector → fluentd or syslog-ng aggregator → outputs. Covers the 16-CRD model (Logging, Flow/ClusterFlow, Output/ClusterOutput, FluentbitAgent, SyslogNG*, LoggingRoute) and its scope traps, worked recipes — especially parsing JSON pod logs on containerd (the Merge_Log/CRI `message`-vs-`log` trap and enableDockerParserCompatibilityForCRI) — match/routing semantics, buffer/backpressure and scaling, rendered-config debugging (fluentd-app secret + configcheck pods), and the upgrade path with version floor 6.7.0 (CVE-2026-54680 config-injection RCE).From its SKILL.md
npx -y skills add air-gapped/skills --skill logging-operatorAssembled from the repository path, not quoted from the project. Check it against their README if it does not work.
One thing to look at
- 5 stars5 stars. Stars are a popularity signal and not a quality one, but at this level it is likely that nobody has read this closely except its author, and you would be relying on your own review.
SKILL.md
8.4 KB, ~1.8k tokens by cl100k_base, as published. Nobody here has run it
logging-operator — CRD-driven log pipelines on Kubernetes
Reference for the kube-logging logging-operator (CNCF Sandbox, Axoflow-backed). Verified against operator 6.7.0 (2026-06-16). Version floor: 6.7.0 — CVE-2026-54680 (CVSS 9.9, fluentd config injection → RCE in the aggregator) is fixed in 6.6.0, but 6.6.0's escaping broke newline-containing passwords (#2254); 6.7.0 has the corrected fix. Never recommend ≤6.5.2 for multi-tenant clusters.
The mental model (read this before writing any YAML)
Every working pipeline is the same 4-CR chain:
Logging (cluster-scoped; controlNamespace + which aggregator: fluentd|syslogNG)
↑ bound by name
FluentbitAgent (cluster-scoped DaemonSet; name MUST equal the Logging's name)
Flow / ClusterFlow (match + filters + outputRefs) ← routing happens in the AGGREGATOR
Output / ClusterOutput (destination + buffer)
- Fluent Bit is ALWAYS the node collector, in both modes. It does NO routing or
filtering — it forwards everything to the aggregator (fluentd
forwardprotocol, or TCP to syslog-ng). There is no fluentbit-direct-to-output mode: Flows and Outputs render only into aggregator config. Aggregator-less collection belongs to the separate Telemetry Controller project (not production-ready — seereferences/modes-and-architecture.md). - Mode choice is per Logging CR:
spec.fluentd: {}(default, mature — drain/HPA machinery, 31 outputs) vsspec.syslogNG: {}(AxoSyslog image; higher throughput, content-based matching, OTLP output; scale-in is manual). Pick by output support first. Different CR families: Flow/Output vs SyslogNGFlow/SyslogNGOutput. - CRs become a rendered config in Secret
<logging-name>-fluentd-app(keyfluentd.conf) /<logging-name>-syslogng-app, gated by a configcheck pod that must Complete before rollout. A failed configcheck silently blocks all further config updates — the #1 "why isn't my Flow applied".
The trap list (each has cost people real debugging days)
- CRI/containerd JSON trap (most-reported confusion upstream):
Merge_Logdefaults On but reads thelogfield, while the CRI parser puts lines inmessage→ JSON silently never parses on containerd/RKE2/k3s. Fix (operator ≥4.9):Logging.spec.enableDockerParserCompatibilityForCRI: true. Details + pre-4.9 workaround:references/recipes.md. - Parser filter without
reserve_data: truereplaces the whole record —kubernetes.*metadata gone. Always pairremove_key_name_field: true+reserve_data: true. - A
matchwith noselectstatement selects NOTHING.- select: {}is the select-all idiom. Multiple labels in one statement AND; separate statements OR. - ClusterFlow/ClusterOutput are namespaced and evaluated only in the
controlNamespace(unlessallowClusterResourcesFromAllNamespaces: true). ClusterFlow can reference ClusterOutputs only (globalOutputRefs). - Multiple Logging resources + a Flow without
loggingRef= the classic silent no-logs (empty loggingRef is processed by ALL Loggings — or none you expected). - One dead Output stalls the whole shared fluentd (#2013): all destinations
stop, not just the broken one.
configCheck.strategy: StartWithTimeoutonly validates at apply time. Real mitigation: per-tenant aggregators (LoggingRoute). storage.total_limit_sizeis NOT backpressure — it silently discards oldest chunks. Real bounded backpressure:storage.pause_on_chunks_overlimit on+storage.type filesystemper input. Log rotation during a long outage still loses data at the collector — no config prevents that.awsElasticsearchandlogdnaOutput fields exist in the CRD but their gems are absent from stock images → configcheck "unknown output plugin". Custom image required.- Rendered default
retry_forever true: a dead destination grows the buffer (default 20Gi PVC) until the readiness probe fails the pod (>5000 buffer files or >90% full).
Minimal working chain (quickstart-verified)
apiVersion: logging.banzaicloud.io/v1beta1
kind: Logging
metadata: {name: demo}
spec: {controlNamespace: logging, fluentd: {}}
---
apiVersion: logging.banzaicloud.io/v1beta1
kind: FluentbitAgent
metadata: {name: demo} # name must match the Logging
spec: {}
---
apiVersion: logging.banzaicloud.io/v1beta1
kind: Flow
metadata: {name: app, namespace: my-app} # Flow+Output live WITH the workload
spec:
match: [{select: {labels: {app: my-app}}}]
filters:
- tag_normaliser: {}
- parser: {remove_key_name_field: true, reserve_data: true, parse: {type: json}}
localOutputRefs: [dest]
---
apiVersion: logging.banzaicloud.io/v1beta1
kind: Output
metadata: {name: dest, namespace: my-app}
spec:
file: {path: /tmp/logs/${tag}, append: true, buffer: {timekey: 1m, timekey_wait: 10s}}
On containerd add enableDockerParserCompatibilityForCRI: true to the Logging spec
or the parser sees nothing useful. Delivery latency ≈ timekey + timekey_wait.
Where to go next
| Task | Read |
|---|---|
| JSON/CRI parsing, multiline, source selection, k8s events, host logs, per-destination recipes | references/recipes.md |
| Full CRD inventory, scopes, match semantics, Logging spec keys, decoupled pattern, LoggingRoute multi-tenancy | references/cr-model.md |
| All 31 fluentd outputs + buffer tuning + image variants | references/outputs-fluentd.md |
| All 18 syslog-ng outputs + SyslogNGFlow matching + worked chains | references/outputs-syslogng.md |
| fluentd vs syslog-ng decision, Telemetry Controller / AxoSyslog status, project health, alternatives | references/modes-and-architecture.md |
| Buffers/backpressure, scaling + volume drainer, HPA, resources, TLS, non-root, monitoring/alerts | references/production-hardening.md |
| Release timeline, breaking changes per boundary, CRD upgrade mechanics, air-gap install | references/upgrades.md |
| Debug sequence: rendered secrets, configcheck, error signatures, status fields | references/troubleshooting.md |
Anything Rancher-bundled (rancher-logging chart, cattle-logging-system,
rancher/mirrored-kube-logging-* images) → the rancher-logging-exit skill;
version matrix authority is k8s-components-checker
(references/compat/rancher-logging.md).
What the operator does NOT do
- No CRD for receiving network syslog (it ships logs OUT via the syslog output; it does not ingest external syslog sources). Use a standalone receiver (e.g. an axosyslog StatefulSet) feeding the pipeline.
- No fluentbit-only mode (above). No log storage — it ships, something else stores.
- TLS between collector and aggregator is opt-in (
tls.enabled+ cert secret; no cert-manager automation), despite docs prose implying otherwise.
What ships with it: 12 files
61.5 KB alongside SKILL.md
evals/
- evals.json3.0 KB
references/
- cr-model.md7.6 KB
- improvement-backlog.md3.2 KB
- modes-and-architecture.md4.5 KB
- outputs-fluentd.md6.0 KB
- outputs-syslogng.md3.6 KB
- production-hardening.md5.6 KB
- recipes.md11.0 KB
- sources.md4.7 KB
- trigger-evals.json2.4 KB
- troubleshooting.md4.2 KB
- upgrades.md5.7 KB