Files
telemt/docs/Advanced_settings/HIGH_LOAD.en.md
T
2026-09-27 18:55:31 +03:00

9.7 KiB
Raw Blame History

High-Load Configuration & Tuning Guide

When deploying Telemt under high-traffic load (tens or hundreds of thousands of concurrent connections), the standard OS network stack limits can lead to packet drops, high CPU context switching, and connection failures. This guide covers Linux kernel tuning, hardware configuration, and architecture optimizations required to prepare the server for high-load scenarios.


1. System Limits & File Descriptors

Every TCP connection requires a file descriptor. At 100k connections, standard Linux limits (often 1024 or 65535) will be exhausted immediately.

System-Wide Limits (sysctl)

Increase the global file descriptor limit in /etc/sysctl.conf:

fs.file-max = 2097152
fs.nr_open = 2097152

User-Level Limits (limits.conf)

Edit /etc/security/limits.conf to allow the telemt (or proxy) user to allocate them:

* soft nofile 1048576
* hard nofile 1048576
root soft nofile 1048576
root hard nofile 1048576

Systemd / Docker Overrides

If using Systemd, add to your telemt.service:

[Service]
LimitNOFILE=1048576
LimitNPROC=65535
TasksMax=infinity

If using Docker, configure ulimits in docker-compose.yaml:

services:
  telemt:
    ulimits:
      nofile:
        soft: 1048576
        hard: 1048576

2. Kernel Network Stack Tuning (sysctl)

Create a dedicated file /etc/sysctl.d/99-telemt-highload.conf and apply it via sysctl -p /etc/sysctl.d/99-telemt-highload.conf.

2.1 Connection Queues & SYN Flood Protection

Increase the size of accept queues to absorb sudden connection spikes (bursts) and mitigate SYN floods:

net.core.somaxconn = 65535
net.core.netdev_max_backlog = 65535
net.ipv4.tcp_max_syn_backlog = 65535
net.ipv4.tcp_syncookies = 1

2.2 Port Exhaustion & TIME-WAIT Sockets

High churn rates lead to ephemeral port exhaustion. Expand the range and rapidly recycle closed sockets:

net.ipv4.ip_local_port_range = 10000 65535
net.ipv4.tcp_fin_timeout = 15
net.ipv4.tcp_tw_reuse = 1
net.ipv4.tcp_max_tw_buckets = 2000000

2.3 TCP Keepalive (Aggressive Dead Connection Culling)

By default, Linux keeps silent, dropped connections open for over 2 hours. This consumes memory at scale. The values below start probing after five idle minutes and abandon an unresponsive peer after the subsequent probe budget, roughly 7–8 minutes after it became idle:

net.ipv4.tcp_keepalive_time = 300
net.ipv4.tcp_keepalive_intvl = 30
net.ipv4.tcp_keepalive_probes = 5

2.4 TCP Buffers & Congestion Control

Optimize memory usage per socket and switch to BBR (Bottleneck Bandwidth and Round-trip propagation time) to improve latency on lossy networks:

# Core buffer sizes
net.core.rmem_default = 262144
net.core.wmem_default = 262144
net.core.rmem_max = 16777216
net.core.wmem_max = 16777216
# TCP-specific buffers (min, default, max)
net.ipv4.tcp_rmem = 4096 87380 16777216
net.ipv4.tcp_wmem = 4096 65536 16777216
# Enable BBR
net.core.default_qdisc = fq
net.ipv4.tcp_congestion_control = bbr

3. Conntrack (Netfilter) Tuning

If your server uses iptables, ufw, or firewalld, the Linux kernel tracks every connection state in a table (nf_conntrack). When this table fills up, Linux drops new packets. Check your current limit and usage:

sysctl net.netfilter.nf_conntrack_max
sysctl net.netfilter.nf_conntrack_count

If it gets close to the limit, tune it up, and reduce the time established connections linger in the tracker:

# In /etc/sysctl.d/99-telemt-highload.conf
net.netfilter.nf_conntrack_max = 2097152
# Reduce timeout from default 5 days to 1 hour
net.netfilter.nf_conntrack_tcp_timeout_established = 3600
net.netfilter.nf_conntrack_tcp_timeout_time_wait = 12

Note: Depending on your OS, you may need to run modprobe nf_conntrack before setting these parameters.

When server.conntrack_control.inline_conntrack_control = true and [server.conntrack_control] uses notrack or hybrid, one generation-fenced authority owns only the conntrack-control rules created by Telemt. It applies IPv4 and IPv6 changes, ignores stale or conflicting publications, and retries failed reconciliation after 1, 2, 4, 8, 16, then capped 30-second delays. A partial multi-command failure triggers best-effort restoration; failed rollback leaves the applied firewall state unknown until a later successful reconcile. telemt_conntrack_control_state{flag="rule_apply_ok"} reports whether the desired rule set is effective. With core telemetry enabled, reconcile and rollback attempts use telemt_conntrack_rule_reconcile_total{result="success"|"error"} and telemt_conntrack_rule_rollback_total{result="success"|"error"}. Shutdown performs a bounded 30-second best-effort cleanup of owned rules. Conntrack-control configuration remains restart-only.


4. Multi-Tier Architecture: HAProxy Setup

4.1 Native MTProxy and TLS-front L4 deployment

For massive native MTProxy or TLS-front traffic, an L4 HAProxy can absorb connection spikes before handing TCP streams to Telemt. The following example is not valid for a WEB listener.

HAProxy High-Load haproxy.cfg

global
    # Disable detailed connection logs under load
    log stdout format raw local0 err
    maxconn 250000
    # Tune buffers and socket acceptance
    tune.bufsize 16384
    tune.maxaccept 64
defaults
    log     global
    mode    tcp
    option  clitcpka
    option  srvtcpka
    timeout connect 5s
    timeout client  1h
    timeout server  1h
    # Purge dead peers quickly
    timeout client-fin 10s
    timeout server-fin 10s
frontend proxy_in
    bind *:443
    maxconn 250000
    option tcp-smart-accept
    default_backend telemt_backend
backend telemt_backend
    option tcp-smart-connect
    # Preserve the client IP for Telemt through PROXY v2
    server telemt_core 10.10.10.1:443 maxconn 250000 send-proxy-v2 check inter 5s

Important: Telemt must be configured to process the PROXY protocol on port 443 for this chain to work and preserve client IPs.

4.2 WEB deployment

WEB mode requires an L7 TLS terminator and a private plain HTTP/1.1 Telemt listener with proxy_protocol = false; do not reuse the L4 send-proxy-v2 backend above. Follow the complete WEB proxy guide and preserve Host, the exact path and query, WebSocket Upgrade headers, and a single overwritten X-Forwarded-For value from explicitly trusted terminator CIDRs. Public ALPN must offer h2 for https-lanes and http/1.1 for WebSocket Upgrade.

Set the WEB listener to web_client_ip_source = "x_forwarded_for" and list only the immediate HAProxy addresses in web_trusted_proxy_cidrs. Never trust a client-reachable subnet and never enable PROXY protocol on this listener.

Route the complete vhost to one Telemt process. For prefix cohosting, preserve the configured base_path without rewrite; while migrating it, route both old and new subtrees to Telemt until old process-issued credentials can no longer be used. Multi-process backends require whole-vhost affinity for the bridge root, session creation, recovery, uplink, downlink, DELETE, diagnostics, and WebSocket Upgrade.

For example, this HAProxy fragment routes only one exact host and slash-terminated WEB subtree without changing the request target:

frontend https_in
    mode http
    bind *:443 ssl crt /etc/haproxy/certs/proxy.pem alpn h2,http/1.1
    timeout client 65s
    acl telemt_web_host hdr(host) -i proxy.example.com proxy.example.com:443
    acl telemt_web_path path_beg /telegram/web/
    use_backend telemt_web if telemt_web_host telemt_web_path

backend telemt_web
    mode http
    retries 0
    timeout connect 5s
    timeout server 65s
    http-request set-header Host proxy.example.com
    http-request set-header X-Forwarded-For %[src]
    server telemt_web_1 127.0.0.1:18080 check

Handle the no-slash /telegram/web path outside this backend so the frontend cannot synthesize a redirect alias into the authenticated subtree. Do not add set-path, replace-path, or a path component to the backend server URL. The 65-second values are examples for the defaults; configure client and server timeouts above web.timeouts.long_poll_secs and twice the effective WebSocket liveness interval.

Capacity planning must include both public TLS sockets and private terminator-to-Telemt sockets. Size file descriptors, terminator upstream capacity, web.limits.max_http_connections, handler capacity, concurrent long polls, and WebSocket lanes together; an upstream keepalive pool is not a concurrency limit.


5. Diagnostics & Monitoring

When operating under load, these commands are useful for diagnostics:

  • Listen queue drops: inspect ListenOverflows and ListenDrops in /proc/net/netstat or nstat.
  • Conntrack pressure: inspect nf_conntrack_count, kernel logs, and telemt_conntrack_control_state{flag="rule_apply_ok"}; with core telemetry enabled, also inspect the reconcile/rollback counters above.
  • File descriptor usage: cat /proc/sys/fs/file-nr and the Telemt process limits under /proc/<pid>/limits.
  • Connection states: ss -s; avoid full netstat scans on a busy host.
  • Rate limiter contention: with core telemetry enabled, alert on a positive counter increase or rate, for example increase(telemt_rate_limiter_cas_retry_exhausted_total[5m]) > 0, grouped by scope, direction, and operation. Reserve exhaustion returns a zero grant without classifying it as a configured throttle; refund exhaustion retains the charge. This metric is neither a connection-drop counter nor a policy-throttle counter.
  • WEB: combine terminator telemetry and an external TLS probe with /v1/runtime/web/status, telemt_web_tcp_accept_total{result="accepted"|"error"}, and the other telemt_web_* metrics. Request paths and base_path are intentionally not metric labels.