Container Runtime Security with Falco for Kubernetes

Implementing real-time threat detection in Kubernetes clusters using Falco for syscall-level monitoring and automated incident response

#kubernetes#security#falco#runtime-security
Cover image for the article: Container Runtime Security with Falco for Kubernetes

Introduction

Container runtime security is the practice of detecting and responding to threats in running containers. While image scanning catches known vulnerabilities before deployment, runtime security detects active exploitation, lateral movement, and anomalous behavior in real-time. Falco, the CNCF-graduated runtime security project, provides syscall-level monitoring with customizable detection rules.

In production Kubernetes environments, Falco typically detects 15-25 security-relevant events per day per 100-pod cluster, with detection latency under 500 milliseconds from syscall to alert. This article covers deployment architecture, rule engineering, and integration with incident response workflows.

Falco Architecture

Falco operates by intercepting system calls at the kernel level using either a kernel module or eBPF probe. Every syscall made by every container is evaluated against a rule engine in real-time.

Chart

Detection Pipeline

StageLatencyFunction
Syscall capture< 1μsKernel module or eBPF probe
Event filtering< 10μsPre-filter uninteresting events
Rule evaluation< 100μsMatch against rule conditions
Alert generation< 1msFormat and emit alert
Alert delivery1-50msSend to configured outputs
Total end-to-end< 100msFrom syscall to alert delivery

Deployment with Helm

# Add Falco Helm repo
helm repo add falcosecurity https://falcosecurity.github.io/charts
helm repo update

# Install Falco with eBPF driver
helm install falco falcosecurity/falco \
  --namespace falco-system \
  --create-namespace \
  --set driver.kind=ebpf \
  --set falcosidekick.enabled=true \
  --set falcosidekick.config.slack.webhookurl="https://hooks.slack.com/services/xxx" \
  --set metrics.enabled=true \
  --set falco.grpc.enabled=true \
  --set falco.grpc_output.enabled=true \
  --values custom-values.yaml

Custom Values Configuration

# custom-values.yaml
falco:
  rules_file:
    - /etc/falco/falco_rules.yaml
    - /etc/falco/falco_rules.local.yaml
    - /etc/falco/rules.d

  json_output: true
  json_include_output_property: true
  log_stderr: true
  log_syslog: false
  log_level: info

  priority: notice

  buffered_outputs: false
  outputs_queue:
    capacity: 1024

  syscall_event_drops:
    actions:
      - log
      - alert
    rate: 0.03333
    max_burst: 1

  grpc:
    enabled: true
    bind_address: "unix:///run/falco/falco.sock"
    threadiness: 0

  grpc_output:
    enabled: true

driver:
  kind: ebpf
  ebpf:
    hostNetwork: true
    leastPrivileged: false

resources:
  requests:
    cpu: 100m
    memory: 512Mi
  limits:
    cpu: 1000m
    memory: 1024Mi

tolerations:
  - effect: NoSchedule
    operator: Exists

Custom Rules for Production

Detecting Container Escape Attempts

- rule: Container Escape via Mount Namespace
  desc: Detect attempts to escape container via mount namespace manipulation
  condition: >
    evt.type in (setns, unshare) and
    container.id != host and
    evt.arg.flags contains CLONE_NEWNS
  output: >
    Container escape attempt via namespace manipulation
    (user=%user.name command=%proc.cmdline container=%container.name
     namespace=%k8s.ns.name pod=%k8s.pod.name image=%container.image.repository)
  priority: CRITICAL
  tags: [container, escape, namespace]

- rule: Suspicious Privilege Escalation in Container
  desc: Detect setuid/setgid or capability manipulation
  condition: >
    spawned_process and
    container and
    (proc.name in (sudo, su, doas) or
     evt.type in (setuid, setgid, capset)) and
    not proc.pname in (entrypoint.sh, docker-entrypoint)
  output: >
    Privilege escalation detected
    (user=%user.name command=%proc.cmdline container=%container.name
     parent=%proc.pname ns=%k8s.ns.name)
  priority: CRITICAL
  tags: [container, privilege_escalation]

Detecting Cryptomining

- rule: Detect Cryptomining Activity
  desc: Detect processes connecting to known mining pools or using mining binaries
  condition: >
    spawned_process and container and
    (proc.name in (xmrig, minerd, minergate, cgminer, bfgminer, cpuminer) or
     proc.cmdline contains "stratum+tcp://" or
     proc.cmdline contains "stratum+ssl://" or
     proc.cmdline contains "--coin" or
     (proc.cmdline contains "-o " and proc.cmdline contains ":3333"))
  output: >
    Cryptomining activity detected
    (command=%proc.cmdline container=%container.name
     image=%container.image.repository ns=%k8s.ns.name pod=%k8s.pod.name)
  priority: CRITICAL
  tags: [container, cryptomining, malware]

- rule: Outbound Connection to Mining Pool
  desc: Detect network connections to known mining pool ports
  condition: >
    outbound and container and
    fd.sport in (3333, 4444, 5555, 7777, 8888, 9999, 14433, 14444)
  output: >
    Outbound connection to potential mining pool
    (command=%proc.cmdline connection=%fd.name container=%container.name
     ns=%k8s.ns.name image=%container.image.repository)
  priority: HIGH
  tags: [container, cryptomining, network]

Detecting Lateral Movement

- rule: Kubernetes Service Account Token Access
  desc: Detect reading of service account tokens from within containers
  condition: >
    open_read and container and
    fd.name startswith /var/run/secrets/kubernetes.io/serviceaccount and
    proc.name != pause and
    not k8s_allowed_sa_readers
  output: >
    Service account token read from container
    (file=%fd.name command=%proc.cmdline container=%container.name
     ns=%k8s.ns.name pod=%k8s.pod.name)
  priority: WARNING
  tags: [kubernetes, lateral_movement, credential_access]

- list: k8s_allowed_sa_readers
  items: [kube-proxy, coredns, calico-node, falco]

Performance Impact

Falco's performance overhead varies by driver type and workload:

DriverCPU OverheadMemorySyscall LatencyPacket Drop Rate
Kernel Module1-3%256-512 MB< 1μs< 0.01%
eBPF (modern)2-5%256-512 MB1-2μs< 0.05%
eBPF (legacy)3-8%512 MB-1 GB2-5μs0.1-0.5%
Plugin (no driver)< 1%128-256 MBN/AN/A

Benchmark: Rule Count vs. Performance

Active RulesEvents/sec (capacity)CPU UsageAlert Latency
50 (default)250,0002%< 50ms
100200,0003%< 75ms
200150,0005%< 100ms
50080,00010%< 200ms

Integration with Falcosidekick

Falcosidekick provides 60+ output destinations for Falco alerts:

# falcosidekick configuration
config:
  slack:
    webhookurl: "https://hooks.slack.com/services/xxx"
    minimumpriority: "warning"
    messageformat: |
      *Priority:* {{ .Priority }}
      *Rule:* {{ .Rule }}
      *Output:* {{ .Output }}
      *Namespace:* {{ index .OutputFields "k8s.ns.name" }}
      *Pod:* {{ index .OutputFields "k8s.pod.name" }}

  pagerduty:
    routingkey: "your-routing-key"
    minimumpriority: "critical"

  prometheus:
    enabled: true
    listen_address: ":9090"

  kuberneteseventgenerator:
    enabled: true
    minimumpriority: "warning"

Automated Response with Falco Talon

# falco-talon rules for automated response
- action: Terminate Pod
  description: Kill pods with critical security violations
  match:
    rules:
      - Container Escape via Mount Namespace
      - Detect Cryptomining Activity
    priority: critical
  responses:
    - action: kubernetes:terminate
      parameters:
        grace_period_seconds: 0

- action: Network Isolation
  description: Isolate pods with suspicious network activity
  match:
    rules:
      - Outbound Connection to Mining Pool
    priority: high
  responses:
    - action: kubernetes:networkpolicy
      parameters:
        allow:
          - "192.168.0.0/16"

Key Takeaways

  • Falco detects threats at the syscall level with under 100ms end-to-end latency from system call to alert delivery, enabling real-time response.
  • Use eBPF driver on modern kernels (5.8+) for the best balance of performance and safety, with 2-5% CPU overhead per node.
  • Custom rules should target your specific threat model rather than relying solely on default rules which generate noise for generic container workloads.
  • Performance scales with rule count so keep active rules under 200 for high-throughput environments processing more than 150,000 events per second.
  • Integrate Falcosidekick for multi-channel alerting routing critical alerts to PagerDuty while sending warnings to Slack for review.
  • Automated pod termination for critical threats like cryptomining and container escapes reduces mean time to containment from minutes to seconds.
  • Monitor Falco itself for dropped events, which indicate the rule engine cannot keep up with syscall volume and requires tuning or scaling.

Comments

    No comments yet. Be the first to share your thoughts.