Howdy, Stranger!

It looks like you're new here. If you want to get involved, click one of these buttons!


New on LowEndTalk? Please Register and read our Community Rules.

All new Registrations are manually reviewed and approved, so a short delay after registration may occur before your account becomes active.

linux-tcp-cc: run upstream Linux BBR on OpenVZ / constrained VPS without changing the host kernel

linux-tcp-cc: run upstream Linux BBR on OpenVZ / constrained VPS without changing the host kernel

Hi,

I've been working on a small open-source project called linux-tcp-cc:

https://github.com/chenmiaoming/linux-tcp-cc

Documentation:

https://tcpcc.250328.xyz/

The goal is fairly specific:

Run the actual upstream Linux TCP stack, including BBR/CUBIC, for public TCP connections on VPS/container environments where you cannot replace the host kernel or change its congestion-control configuration.

The main use case is old-style OpenVZ / constrained containers where you may have TUN and netfilter access, but the provider kernel owns normal TCP sockets.

Why I made this

This is not a new idea.

Years ago, projects such as lkl-haproxy used Linux Kernel Library + HAProxy + TAP/iptables to get BBR/BBRPlus working inside OpenVZ.

There are still maintained variants of that idea today, for example LKL builds using LD_PRELOAD=liblkl-hijack.so.

I liked the fundamental idea:

If the host kernel owns the public TCP socket, you cannot change its congestion control. So move ownership of that TCP socket somewhere else.

What I wanted to change was the deployment model.

Instead of:

patched LKL
    +
liblkl-hijack.so
    +
LD_PRELOAD
    +
specific HAProxy compatibility
    +
TAP setup
    +
iptables scripts
    +
systemd glue

I wanted something closer to a normal standalone networking tool:

sudo tcpcc \
  --forward 203.0.113.10:443=127.0.0.1:8443 \
  --cc bbr

or:

version = 1
cc = "bbr"
memory_mib = 128

[[forward]]
listen = "203.0.113.10:443"
backend = "127.0.0.1:8443"

Then:

sudo tcpcc --config /etc/tcpcc/tcpcc.toml

The important design goal: upstream Linux fidelity

The main reason this project exists is not because writing another userspace TCP stack is impossible.

The point is the opposite: I specifically do not want to reimplement Linux TCP behavior.

The public connection is terminated inside a hosted upstream Linux networking stack:

remote client
    |
    | public TCP packets
    v
host DNAT / conntrack
    |
    v
TUN
    |
    v
hosted upstream Linux TCP listener
    |  TCP_CONGESTION = bbr / cubic
    |  upstream Linux recovery
    |  upstream rate sampling
    |  upstream fq pacing
    v
byte-stream bridge
    |
    v
127.0.0.1 backend

The backend can still be a completely normal nginx, HAProxy, web server, etc.

It does not need to run under LD_PRELOAD.

This is one of the main design goals of linux-tcp-cc:

keep the TCP behavior as close to upstream Linux as possible.

The project deliberately does not reimplement BBR, CUBIC, Linux delivery-rate sampling, TCP loss recovery, or fq pacing.

Those are the real upstream Linux implementations running inside the hosted stack.

For this project, using upstream Linux is not just an implementation shortcut. It is the feature.

The current branch tracks Linux 6.18.y.

Why not just implement a smaller userspace TCP stack?

That is actually another direction I am working on.

I have another project called tcp-shift:

https://github.com/chenmiaoming/tcp-shift

It uses lwIP and targets a different point in the design space.

tcp-shift is intended to be a much smaller userspace TCP endpoint, with goals such as:

  • much lower fixed memory overhead;
  • compact per-flow state;
  • explicit RFC-oriented transport behavior;
  • Reno/CUBIC/BBR policy integration;
  • RFC 8985 RACK-TLP recovery;
  • event-driven operation without busy polling.

For example, its current warm process PSS is around:

507 KiB

before flow state.

But the tradeoff is important:

tcp-shift necessarily owns and validates more TCP semantics itself.

So I see these as two complementary projects:

linux-tcp-cc
    |
    +-- priority: upstream Linux fidelity
    +-- actual upstream Linux TCP implementation
    +-- BBR/CUBIC/recovery/rate sampling/fq come from Linux
    +-- larger runtime footprint

tcp-shift
    |
    +-- priority: small memory footprint
    +-- lwIP transport
    +-- RFC-oriented implementation/qualification
    +-- project-owned congestion-control integration

If I want to experiment with a compact TCP implementation and RFC behavior, that work belongs in tcp-shift.

For linux-tcp-cc, I intentionally chose the other side of the tradeoff:

run upstream Linux networking code and avoid subtly reproducing Linux TCP semantics in another implementation.

IPv4 and IPv6

Public listeners can be IPv4 or IPv6.

For example:

sudo tcpcc \
  --forward '[2001:db8::10]:443=127.0.0.1:8443' \
  --cc bbr

The public IPv6 TCP connection terminates in hosted Linux and can still bridge to an ordinary IPv4 loopback backend.

Multiple forwards are also supported:

sudo tcpcc \
  --forward 203.0.113.10:443=127.0.0.1:8443 \
  --forward 203.0.113.10:8443=127.0.0.1:9443 \
  --cc bbr

Currently, public listeners within one process use the same address family.

Resource usage

I've also been trying to keep the hosted-kernel approach usable on actual low-memory VPSes rather than treating it only as an architecture experiment.

I tested it on a real 128 MiB OpenVZ VPS.

In that field test, the hosted Linux process was roughly:

idle RSS:             ~6.1 MiB
observed peak RSS:   ~17.6 MiB
after flow:           ~6.2 MiB
supervisor RSS:       ~2 MiB

This is a field observation, not a universal memory guarantee.

The hosted memory arena is demand-backed, so configuring:

memory_mib = 128

does not mean 128 MiB immediately becomes resident on the host.

There is also controlled CI qualification for memory capacity, reclaim and reuse, including a 16,384-connection run without restarting the hosted kernel.

If absolute memory footprint is the primary goal, tcp-shift is probably the more interesting direction.

If upstream Linux TCP behavior is the primary goal, that is what linux-tcp-cc is designed for.

Network testing

I also added controlled network qualification rather than only testing on my own VPS.

One deliberately nasty scenario uses:

1 Gbit/s link
200 ms RTT
10% random loss in each direction
single TCP stream
3 repetitions

The important part for me is not claiming some universal "X times faster" number from one artificial test.

The purpose is to verify things such as:

  • the requested congestion control really belongs to the hosted public listener;
  • real impaired traffic passes through the hosted Linux TCP stack;
  • the selected BBR/CUBIC implementation is the Linux implementation;
  • connections survive the complete packet path;
  • shutdown and resource cleanup work;
  • memory can be reclaimed and the same hosted process reused;
  • the runtime does not require periodic polling to stay alive.

Reviewed evidence is kept in the main repository under:

benchmarks/published/

Documentation:

https://tcpcc.250328.xyz/benchmarks/

Requirements / limitations

This is not magic and it cannot bypass every OpenVZ restriction.

You still need:

  • /dev/net/tun;
  • sufficient network privileges to configure the TUN path;
  • nftables or iptables;
  • IP forwarding;
  • a provider that does not completely prohibit this packet path.

The backend is currently:

127.0.0.1:<port>

The public connection and backend connection are deliberately two separate TCP connections.

--cc bbr applies to the hosted public TCP listener.

It does not change the outer host kernel's loopback TCP congestion control.

Related / prior work

The closest prior art I found is the LKL + HAProxy family:

  • tcp-nanqinlang/lkl-haproxy
  • mzz2017/lkl-haproxy
  • nivrrex/lkl-bbr
  • nivrrex/lkl-proxy

Those projects are the reason I would not claim that "BBR on OpenVZ from userspace" is a new concept.

In fact, they demonstrated the core idea years ago.

What I'm trying to contribute is a different maintenance and application boundary.

Roughly:

traditional LKL-haproxy approach

application / proxy
        |
        | LD_PRELOAD syscall hijacking
        v
       LKL
        |
       TAP

versus:

linux-tcp-cc

remote TCP
    |
   TUN
    |
hosted upstream Linux TCP
    |
byte-stream bridge
    |
ordinary 127.0.0.1 application

The backend application does not have to know that the hosted Linux stack exists.

And most importantly, the design goal is to preserve upstream Linux TCP behavior, rather than building an approximation of Linux BBR around another TCP implementation.

Feedback wanted

This is still a fairly niche project, so LowEndTalk seems like one of the few places where people may actually have exactly the kind of terrible VPS this is intended for :)

I'm particularly interested in testing on:

  • old OpenVZ / Virtuozzo environments;
  • VPSes where BBR is unavailable but TUN works;
  • 64 MiB / 128 MiB low-memory VPSes;
  • IPv6-only environments;
  • NAT-heavy setups;
  • long-RTT or lossy international routes;
  • strange nftables/iptables environments.

If anyone here still runs one of the old LKL + HAProxy BBR setups, I'd also be very interested in comparisons.

In particular, I'd like to know which operational problems from the older approach still matter in real deployments and which ones I may have missed.

Source:

https://github.com/chenmiaoming/linux-tcp-cc

Documentation:

https://tcpcc.250328.xyz/

Downloads:

https://github.com/chenmiaoming/linux-tcp-cc/releases

License:

GPL-2.0

Comments

  • forestforest Member

    Haven't read the whole slopwall yet but why OpenVZ? Why not LXC?

  • ObelousObelous Member

    Please leave your slop at the door, thank you!

  • miaomingcmiaomingc Member
    edited 4:34PM

    @forest said:
    Haven't read the whole slopwall yet but why OpenVZ? Why not LXC?

    LXC itself doesn't give the container control over the host kernel either.

    What I meant is that if you are the LXC host administrator, you can normally enable BBR in the host kernel and there is little reason to use tcpcc.

    For a rented LXC container, the situation can be very similar to OpenVZ: if the provider kernel does not provide BBR, or the container is not allowed to select it, the guest cannot add BBR itself.

    So tcpcc is not OpenVZ-specific. OpenVZ is just the motivating case because this limitation is especially common there. The actual target is any constrained container/VPS where TUN/network privileges are available but the desired TCP CC is not.

  • yoursunnyyoursunny Member, IPv6 Advocate

    Does this work on VZ6, asking for a friend?

Sign In or Register to comment.