Eric Lee / smarc-fsl-linux-kernel

23 Sep, 2016

1 commit

fefa569a9 net_sched: sch_fq: account for schedule/timers drifts ... Browse Code »

It looks like the following patch can make FQ very precise, even in VM
or stressed hosts. It matters at high pacing rates.

We take into account the difference between the time that was programmed
when last packet was sent, and current time (a drift of tens of usecs is
often observed)

Add an EWMA of the unthrottle latency to help diagnostics.

This latency is the difference between current time and oldest packet in
delayed RB-tree. This accounts for the high resolution timer latency,
but can be different under stress, as fq_check_throttled() can be
opportunistically be called from a dequeue() called after an enqueue()
for a different flow.

Tested:
// Start a 10Gbit flow
$ netperf --google-pacing-rate 1250000000 -H lpaa24 -l 10000 -- -K bbr &

Before patch :
$ sar -n DEV 10 5 | grep eth0 | grep Average
Average: eth0 17106.04 756876.84 1102.75 1119049.02 0.00 0.00 0.52

After patch :
$ sar -n DEV 10 5 | grep eth0 | grep Average
Average: eth0 17867.00 800245.90 1151.77 1183172.12 0.00 0.00 0.52

A new iproute2 tc can output the 'unthrottle latency' :

$ tc -s qd sh dev eth0 | grep latency
0 gc, 0 highprio, 32490767 throttled, 2382 ns latency

Signed-off-by: Eric Dumazet
Signed-off-by: David S. Miller

Eric Dumazet
2016-09-23 19:19:06 +0800

21 Sep, 2016

1 commit

77879147a net_sched: sch_fq: add low_rate_threshold parameter ... Browse Code »

This commit adds to the fq module a low_rate_threshold parameter to
insert a delay after all packets if the socket requests a pacing rate
below the threshold.

This helps achieve more precise control of the sending rate with
low-rate paths, especially policers. The basic issue is that if a
congestion control module detects a policer at a certain rate, it may
want fq to be able to shape to that policed rate. That way the sender
can avoid policer drops by having the packets arrive at the policer at
or just under the policed rate.

The default threshold of 550Kbps was chosen analytically so that for
policers or links at 500Kbps or 512Kbps fq would very likely invoke
this mechanism, even if the pacing rate was briefly slightly above the
available bandwidth. This value was then empirically validated with
two years of production testing on YouTube video servers.

Signed-off-by: Van Jacobson
Signed-off-by: Neal Cardwell
Signed-off-by: Yuchung Cheng
Signed-off-by: Nandita Dukkipati
Signed-off-by: Eric Dumazet
Signed-off-by: Soheil Hassas Yeganeh
Signed-off-by: David S. Miller

Eric Dumazet
2016-09-21 12:23:00 +0800

19 Sep, 2016

1 commit

695b4ec0f pkt_sched: fq: use proper locking in fq_dump_stats() ... Browse Code »

When fq is used on 32bit kernels, we need to lock the qdisc before
copying 64bit fields.

Otherwise "tc -s qdisc ..." might report bogus values.

Fixes: afe4fd062416 ("pkt_sched: fq: Fair Queue packet scheduler")
Signed-off-by: Eric Dumazet
Signed-off-by: David S. Miller

Eric Dumazet
2016-09-19 10:15:08 +0800

26 Jun, 2016

1 commit

520ac30f4 net_sched: drop packets after root qdisc lock is released ... Browse Code »

Qdisc performance suffers when packets are dropped at enqueue()
time because drops (kfree_skb()) are done while qdisc lock is held,
delaying a dequeue() draining the queue.

Nominal throughput can be reduced by 50 % when this happens,
at a time we would like the dequeue() to proceed as fast as possible.

Even FQ is vulnerable to this problem, while one of FQ goals was
to provide some flow isolation.

This patch adds a 'struct sk_buff **to_free' parameter to all
qdisc->enqueue(), and in qdisc_drop() helper.

I measured a performance increase of up to 12 %, but this patch
is a prereq so that future batches in enqueue() can fly.

Signed-off-by: Eric Dumazet
Acked-by: Jesper Dangaard Brouer
Signed-off-by: David S. Miller

Eric Dumazet
2016-06-26 00:19:35 +0800

16 Jun, 2016

1 commit

e14ffdfdd net_sched: sch_fq: defer skb freeing ... Browse Code »

Both fq_change() and fq_reset() can use rtnl_kfree_skbs()

Signed-off-by: Eric Dumazet
Signed-off-by: David S. Miller

Eric Dumazet
2016-06-16 05:08:35 +0800

11 Jun, 2016

1 commit

45f50bed1 net_sched: remove generic throttled management ... Browse Code »

__QDISC_STATE_THROTTLED bit manipulation is rather expensive
for HTB and few others.

I already removed it for sch_fq in commit f2600cf02b5b
("net: sched: avoid costly atomic operation in fq_dequeue()")
and so far nobody complained.

When one ore more packets are stuck in one or more throttled
HTB class, a htb dequeue() performs two atomic operations
to clear/set __QDISC_STATE_THROTTLED bit, while root qdisc
lock is held.

Removing this pair of atomic operations bring me a 8 % performance
increase on 200 TCP_RR tests, in presence of throttled classes.

This patch has no side effect, since nothing actually uses
disc_is_throttled() anymore.

Signed-off-by: Eric Dumazet
Signed-off-by: David S. Miller

Eric Dumazet
2016-06-11 14:58:21 +0800

01 Mar, 2016

1 commit

2ccccf5fb net_sched: update hierarchical backlog too ... Browse Code »

When the bottom qdisc decides to, for example, drop some packet,
it calls qdisc_tree_decrease_qlen() to update the queue length
for all its ancestors, we need to update the backlog too to
keep the stats on root qdisc accurate.

Cc: Jamal Hadi Salim
Acked-by: Jamal Hadi Salim
Signed-off-by: Cong Wang
Signed-off-by: David S. Miller

WANG Cong
2016-03-01 06:02:33 +0800

11 Oct, 2015

1 commit

e446f9dfe net: synack packets can be attached to request sockets ... Browse Code »

selinux needs few changes to accommodate fact that SYNACK messages
can be attached to a request socket, lacking sk_security pointer

(Only syncookies are still attached to a TCP_LISTEN socket)

Adds a new sk_listener() helper, and use it in selinux and sch_fq

Fixes: ca6fb0651883 ("tcp: attach SYNACK messages to request sockets instead of listener")
Signed-off-by: Eric Dumazet
Reported by: kernel test robot
Cc: Paul Moore
Cc: Stephen Smalley
Cc: Eric Paris
Acked-by: Paul Moore
Signed-off-by: David S. Miller

Eric Dumazet
2015-10-11 20:05:06 +0800

03 Oct, 2015

1 commit

ca6fb0651 tcp: attach SYNACK messages to request sockets instead of listener ... Browse Code »

If a listen backlog is very big (to avoid syncookies), then
the listener sk->sk_wmem_alloc is the main source of false
sharing, as we need to touch it twice per SYNACK re-transmit
and TX completion.

(One SYN packet takes listener lock once, but up to 6 SYNACK
are generated)

By attaching the skb to the request socket, we remove this
source of contention.

Tested:

listen(fd, 10485760); // single listener (no SO_REUSEPORT)
16 RX/TX queue NIC
Sustain a SYNFLOOD attack of ~320,000 SYN per second,
Sending ~1,400,000 SYNACK per second.
Perf profiles now show listener spinlock being next bottleneck.

20.29% [kernel] [k] queued_spin_lock_slowpath
10.06% [kernel] [k] __inet_lookup_established
5.12% [kernel] [k] reqsk_timer_handler
3.22% [kernel] [k] get_next_timer_interrupt
3.00% [kernel] [k] tcp_make_synack
2.77% [kernel] [k] ipt_do_table
2.70% [kernel] [k] run_timer_softirq
2.50% [kernel] [k] ip_finish_output
2.04% [kernel] [k] cascade

Signed-off-by: Eric Dumazet
Signed-off-by: David S. Miller

Eric Dumazet
2015-10-03 19:32:43 +0800

02 Apr, 2015

1 commit

05e8bb860 pkt_sched: fq: correct spelling of locally ... Browse Code »

Correct spelling of locally.

Also remove extra space before tab character in struct fq_flow.

Signed-off-by: Simon Horman
Signed-off-by: David S. Miller

Simon Horman
2015-04-02 10:52:29 +0800

06 Feb, 2015

1 commit

6e03f896b Merge git://git.kernel.org/pub/scm/linux/kernel/git/davem/net ... Browse Code »

Conflicts:
drivers/net/vxlan.c
drivers/vhost/net.c
include/linux/if_vlan.h
net/core/dev.c

The net/core/dev.c conflict was the overlap of one commit marking an
existing function static whilst another was adding a new function.

In the include/linux/if_vlan.h case, the type used for a local
variable was changed in 'net', whereas the function got rewritten
to fix a stacked vlan bug in 'net-next'.

In drivers/vhost/net.c, Al Viro's iov_iter conversions in 'net-next'
overlapped with an endainness fix for VHOST 1.0 in 'net'.

In drivers/net/vxlan.c, vxlan_find_vni() added a 'flags' parameter
in 'net-next' whereas in 'net' there was a bug fix to pass in the
correct network namespace pointer in calls to this function.

Signed-off-by: David S. Miller

David S. Miller
2015-02-06 06:33:28 +0800

05 Feb, 2015

3 commits

06eb395fa pkt_sched: fq: better control of DDOS traffic ... Browse Code »

FQ has a fast path for skb attached to a socket, as it does not
have to compute a flow hash. But for other packets, FQ being non
stochastic means that hosts exposed to random Internet traffic
can allocate million of flows structure (104 bytes each) pretty
easily. Not only host can OOM, but lookup in RB trees can take
too much cpu and memory resources.

This patch adds a new attribute, orphan_mask, that is adding
possibility of having a stochastic hash for orphaned skb.

Its default value is 1024 slots, to mimic SFQ behavior.

Note: This does not apply to locally generated TCP traffic,
and no locally generated traffic will share a flow structure
with another perfect or stochastic flow.

This patch also handles the specific case of SYNACK messages:

They are attached to the listener socket, and therefore all map
to a single hash bucket. If listener have set SO_MAX_PACING_RATE,
hoping to have new accepted socket inherit this rate, SYNACK
might be paced and even dropped.

This is very similar to an internal patch Google have used more
than one year.

Signed-off-by: Eric Dumazet
Signed-off-by: David S. Miller

Eric Dumazet
2015-02-05 14:15:45 +0800
987819657 tcp: do not pace pure ack packets ... Browse Code »

When we added pacing to TCP, we decided to let sch_fq take care
of actual pacing.

All TCP had to do was to compute sk->pacing_rate using simple formula:

sk->pacing_rate = 2 * cwnd * mss / rtt

It works well for senders (bulk flows), but not very well for receivers
or even RPC :

cwnd on the receiver can be less than 10, rtt can be around 100ms, so we
can end up pacing ACK packets, slowing down the sender.

Really, only the sender should pace, according to its own logic.

Instead of adding a new bit in skb, or call yet another flow
dissection, we tweak skb->truesize to a small value (2), and
we instruct sch_fq to use new helper and not pace pure ack.

Note this also helps TCP small queue, as ack packets present
in qdisc/NIC do not prevent sending a data packet (RPC workload)

This helps to reduce tx completion overhead, ack packets can use regular
sock_wfree() instead of tcp_wfree() which is a bit more expensive.

This has no impact in the case packets are sent to loopback interface,
as we do not coalesce ack packets (were we would detect skb->truesize
lie)

In case netem (with a delay) is used, skb_orphan_partial() also sets
skb->truesize to 1.

This patch is a combination of two patches we used for about one year at
Google.

Signed-off-by: Eric Dumazet
Signed-off-by: David S. Miller

Eric Dumazet
2015-02-05 12:36:31 +0800
3725a2698 pkt_sched: fq: avoid hang when quantum 0 ... Browse Code »

Configuring fq with quantum 0 hangs the system, presumably because of a
non-interruptible infinite loop. Either way quantum 0 does not make sense.

Reproduce with:
sudo tc qdisc add dev lo root fq quantum 0 initial_quantum 0
ping 127.0.0.1

Signed-off-by: Kenneth Klette Jonassen
Acked-by: Eric Dumazet
Signed-off-by: David S. Miller

Kenneth Klette Jonassen
2015-02-05 12:07:39 +0800

29 Jan, 2015

1 commit

86b3bfe91 pkt_sched: fq: remove useless TIME_WAIT check ... Browse Code »

TIME_WAIT sockets are not owning any skb.

ip_send_unicast_reply() and tcp_v6_send_response() both use
regular sockets.

We can safely remove a test in sch_fq and save one cache line miss,
as sk_state is far away from sk_pacing_rate.

Tested at Google for about one year.

Signed-off-by: Eric Dumazet
Signed-off-by: David S. Miller

Eric Dumazet
2015-01-29 15:23:57 +0800

27 Nov, 2014

1 commit

ced7a04e3 pkt_sched: fq: increase max delay from 125 ms to one second ... Browse Code »

FQ/pacing has a clamp of delay of 125 ms, to avoid some possible harm.

It turns out this delay is too small to allow pacing low rates :
Some ISP setup very aggressive policers as low as 16kbit.

Now TCP stack has spurious rtx prevention, it seems safe to increase
this fixed parameter, without adding a qdisc attribute.

Signed-off-by: Eric Dumazet
Cc: Yang Yingliang
Signed-off-by: David S. Miller

Eric Dumazet
2014-11-27 01:08:04 +0800

06 Oct, 2014

1 commit

f2600cf02 net: sched: avoid costly atomic operation in fq_dequeue() ... Browse Code »

Standard qdisc API to setup a timer implies an atomic operation on every
packet dequeue : qdisc_unthrottled()

It turns out this is not really needed for FQ, as FQ has no concept of
global qdisc throttling, being a qdisc handling many different flows,
some of them can be throttled, while others are not.

Fix is straightforward : add a 'bool throttle' to
qdisc_watchdog_schedule_ns(), and remove calls to qdisc_unthrottled()
in sch_fq.

Signed-off-by: Eric Dumazet
Signed-off-by: David S. Miller

Eric Dumazet
2014-10-06 12:55:10 +0800

30 Sep, 2014

1 commit

25331d6ce net: sched: implement qstat helper routines ... Browse Code »

This adds helpers to manipulate qstats logic and replaces locations
that touch the counters directly. This simplifies future patches
to push qstats onto per cpu counters.

Signed-off-by: John Fastabend
Signed-off-by: David S. Miller

John Fastabend
2014-09-30 13:02:26 +0800

23 Aug, 2014

1 commit

d2de875c6 net: use ktime_get_ns() and ktime_get_real_ns() helpers ... Browse Code »

ktime_get_ns() replaces ktime_to_ns(ktime_get())

ktime_get_real_ns() replaces ktime_to_ns(ktime_get_real())

Signed-off-by: Eric Dumazet
Signed-off-by: David S. Miller

Eric Dumazet
2014-08-23 10:57:23 +0800

05 Jun, 2014

1 commit

4cb28970a net: use the new API kvfree() ... Browse Code »

It is available since v3.15-rc5.

Cc: Pablo Neira Ayuso
Cc: "David S. Miller"
Signed-off-by: Cong Wang
Signed-off-by: David S. Miller

WANG Cong
2014-06-05 15:49:51 +0800

14 Mar, 2014

1 commit

d59b7d805 net_sched: return nla_nest_end() instead of skb->len ... Browse Code »

nla_nest_end() already has return skb->len, so replace
return skb->len with return nla_nest_end instead().

Signed-off-by: Yang Yingliang
Signed-off-by: David S. Miller

Yang Yingliang
2014-03-14 03:39:20 +0800

09 Mar, 2014

1 commit

2d8d40afd pkt_sched: fq: do not hold qdisc lock while allocating memory ... Browse Code »

Resizing fq hash table allocates memory while holding qdisc spinlock,
with BH disabled.

This is definitely not good, as allocation might sleep.

We can drop the lock and get it when needed, we hold RTNL so no other
changes can happen at the same time.

Signed-off-by: Eric Dumazet
Fixes: afe4fd062416 ("pkt_sched: fq: Fair Queue packet scheduler")
Signed-off-by: David S. Miller

Eric Dumazet
2014-03-09 08:09:10 +0800

18 Dec, 2013

2 commits

3958afa1b net: Change skb_get_rxhash to skb_get_hash ... Browse Code »

Changing name of function as part of making the hash in skbuff to be
generic property, not just for receive path.

Signed-off-by: Tom Herbert
Signed-off-by: David S. Miller

Tom Herbert
2013-12-18 05:36:21 +0800
c3bd85495 pkt_sched: fq: more robust memory allocation ... Browse Code »

This patch brings NUMA support and automatic fallback to vmalloc()
in case kmalloc() failed to allocate FQ hash table.

NUMA support depends on XPS being setup for the device before
qdisc allocation. After a XPS change, it might be worth creating
qdisc hierarchy again.

Signed-off-by: Eric Dumazet
Signed-off-by: David S. Miller

Eric Dumazet
2013-12-18 04:25:20 +0800

16 Nov, 2013

2 commits

f52ed8997 pkt_sched: fq: fix pacing for small frames ... Browse Code »

For performance reasons, sch_fq tried hard to not setup timers for every
sent packet, using a quantum based heuristic : A delay is setup only if
the flow exhausted its credit.

Problem is that application limited flows can refill their credit
for every queued packet, and they can evade pacing.

This problem can also be triggered when TCP flows use small MSS values,
as TSO auto sizing builds packets that are smaller than the default fq
quantum (3028 bytes)

This patch adds a 40 ms delay to guard flow credit refill.

Fixes: afe4fd062416 ("pkt_sched: fq: Fair Queue packet scheduler")
Signed-off-by: Eric Dumazet
Cc: Maciej Żenczykowski
Cc: Willem de Bruijn
Cc: Yuchung Cheng
Cc: Neal Cardwell
Signed-off-by: David S. Miller

Eric Dumazet
2013-11-16 10:01:52 +0800
65c5189a2 pkt_sched: fq: warn users using defrate ... Browse Code »

Commit 7eec4174ff29 ("pkt_sched: fq: fix non TCP flows pacing")
obsoleted TCA_FQ_FLOW_DEFAULT_RATE without notice for the users.

Suggested by David Miller

Signed-off-by: Eric Dumazet
Signed-off-by: David S. Miller

Eric Dumazet
2013-11-16 10:01:52 +0800

15 Nov, 2013

1 commit

2abc2f070 pkt_sched: fq: change classification of control packets ... Browse Code »

Initial sch_fq implementation copied code from pfifo_fast to classify
a packet as a high prio packet.

This clashes with setups using PRIO with say 7 bands, as one of the
band could be incorrectly (mis)classified by FQ.

Packets would be queued in the 'internal' queue, and no pacing ever
happen for this special queue.

Fixes: afe4fd062416 ("pkt_sched: fq: Fair Queue packet scheduler")
Signed-off-by: Maciej Żenczykowski
Signed-off-by: Eric Dumazet
Cc: Stephen Hemminger
Cc: Willem de Bruijn
Cc: Yuchung Cheng
Signed-off-by: David S. Miller

Maciej Żenczykowski
2013-11-15 06:16:07 +0800

28 Oct, 2013

1 commit

fc59d5bdf pkt_sched: fq: clear time_next_packet for reused flows ... Browse Code »

When a socket is freed/reallocated, we need to clear time_next_packet
or else we can inherit a prior value and delay first packets of the
new flow.

Signed-off-by: Eric Dumazet
Signed-off-by: David S. Miller

Eric Dumazet
2013-10-28 12:18:31 +0800

09 Oct, 2013

2 commits

7eec4174f pkt_sched: fq: fix non TCP flows pacing ... Browse Code »

Steinar reported FQ pacing was not working for UDP flows.

It looks like the initial sk->sk_pacing_rate value of 0 was
a wrong choice. We should init it to ~0U (unlimited)

Then, TCA_FQ_FLOW_DEFAULT_RATE should be removed because it makes
no real sense. The default rate is really unlimited, and we
need to avoid a zero divide.

Reported-by: Steinar H. Gunderson
Signed-off-by: Eric Dumazet
Signed-off-by: David S. Miller

Eric Dumazet
2013-10-09 09:54:01 +0800
ede869cd0 pkt_sched: fq: fix typo for initial_quantum ... Browse Code »

TCA_FQ_INITIAL_QUANTUM should set q->initial_quantum

Signed-off-by: Eric Dumazet
Signed-off-by: David S. Miller

Eric Dumazet
2013-10-09 04:32:41 +0800

02 Oct, 2013

1 commit

0eab5eb7a pkt_sched: fq: rate limiting improvements ... Browse Code »

FQ rate limiting suffers from two problems, reported
by Steinar :

1) FQ enforces a delay when flow quantum is exhausted in order
to reduce cpu overhead. But if packets are small, current
delay computation is slightly wrong, and observed rates can
be too high.

Steinar had this problem because he disabled TSO and GSO,
and default FQ quantum is 2*1514.

(Of course, I wish recent TSO auto sizing changes will help
to not having to disable TSO in the first place)

2) maxrate was not used for forwarded flows (skbs not attached
to a socket)

Tested:

tc qdisc add dev eth0 root est 1sec 4sec fq maxrate 8Mbit
netperf -H lpq84 -l 1000 &
sleep 10 ; tc -s qdisc show dev eth0
qdisc fq 8003: root refcnt 32 limit 10000p flow_limit 100p buckets 1024
quantum 3028 initial_quantum 15140 maxrate 8000Kbit
Sent 16819357 bytes 11258 pkt (dropped 0, overlimits 0 requeues 0)
rate 7831Kbit 653pps backlog 7570b 5p requeues 0
44 flows (43 inactive, 1 throttled), next packet delay 2977352 ns
0 gc, 0 highprio, 5545 throttled

lpq83:~# tcpdump -p -i eth0 host lpq84 -c 12
09:02:52.079484 IP lpq83 > lpq84: . 1389536928:1389538376(1448) ack 3808678021 win 457
09:02:52.079499 IP lpq83 > lpq84: . 1448:2896(1448) ack 1 win 457
09:02:52.079906 IP lpq84 > lpq83: . ack 2896 win 16384
09:02:52.082568 IP lpq83 > lpq84: . 2896:4344(1448) ack 1 win 457
09:02:52.082581 IP lpq83 > lpq84: . 4344:5792(1448) ack 1 win 457
09:02:52.083017 IP lpq84 > lpq83: . ack 5792 win 16384
09:02:52.085678 IP lpq83 > lpq84: . 5792:7240(1448) ack 1 win 457
09:02:52.085693 IP lpq83 > lpq84: . 7240:8688(1448) ack 1 win 457
09:02:52.086117 IP lpq84 > lpq83: . ack 8688 win 16384
09:02:52.088792 IP lpq83 > lpq84: . 8688:10136(1448) ack 1 win 457
09:02:52.088806 IP lpq83 > lpq84: . 10136:11584(1448) ack 1 win 457
09:02:52.089217 IP lpq84 > lpq83: . ack 11584 win 16384

Reported-by: Steinar H. Gunderson
Signed-off-by: Eric Dumazet
Signed-off-by: David S. Miller

Eric Dumazet
2013-10-02 01:00:38 +0800

01 Oct, 2013

1 commit

8d34ce10c pkt_sched: fq: qdisc dismantle fixes ... Browse Code »

fq_reset() should drops all packets in queue, including
throttled flows.

This patch moves code from fq_destroy() to fq_reset()
to do the cleaning.

fq_change() must stop calling fq_dequeue() if all remaining
packets are from throttled flows.

Signed-off-by: Eric Dumazet
Signed-off-by: David S. Miller

Eric Dumazet
2013-10-01 03:51:23 +0800

31 Aug, 2013

1 commit

08f89b981 pkt_sched: fq: prefetch() fix ... Browse Code »

kbuild bot reported following m68k build error :

net/sched/sch_fq.c: In function 'fq_dequeue':
>> net/sched/sch_fq.c:491:2: error: implicit declaration of function
'prefetch' [-Werror=implicit-function-declaration]
cc1: some warnings being treated as errors

While we are fixing this, move this prefetch() call a bit earlier.

Reported-by: Wu Fengguang
Signed-off-by: Eric Dumazet
Signed-off-by: David S. Miller

Eric Dumazet
2013-08-31 02:51:59 +0800

30 Aug, 2013

1 commit

afe4fd062 pkt_sched: fq: Fair Queue packet scheduler ... Browse Code »

- Uses perfect flow match (not stochastic hash like SFQ/FQ_codel)
- Uses the new_flow/old_flow separation from FQ_codel
- New flows get an initial credit allowing IW10 without added delay.
- Special FIFO queue for high prio packets (no need for PRIO + FQ)
- Uses a hash table of RB trees to locate the flows at enqueue() time
- Smart on demand gc (at enqueue() time, RB tree lookup evicts old
unused flows)
- Dynamic memory allocations.
- Designed to allow millions of concurrent flows per Qdisc.
- Small memory footprint : ~8K per Qdisc, and 104 bytes per flow.
- Single high resolution timer for throttled flows (if any).
- One RB tree to link throttled flows.
- Ability to have a max rate per flow. We might add a socket option
to add per socket limitation.

Attempts have been made to add TCP pacing in TCP stack, but this
seems to add complex code to an already complex stack.

TCP pacing is welcomed for flows having idle times, as the cwnd
permits TCP stack to queue a possibly large number of packets.

This removes the 'slow start after idle' choice, hitting badly
large BDP flows, and applications delivering chunks of data
as video streams.

Nicely spaced packets :
Here interface is 10Gbit, but flow bottleneck is ~20Mbit

cwin is big, yet FQ avoids the typical bursts generated by TCP
(as in netperf TCP_RR -- -r 100000,100000)

15:01:23.545279 IP A > B: . 78193:81089(2896) ack 65248 win 3125
15:01:23.545394 IP B > A: . ack 81089 win 3668
15:01:23.546488 IP A > B: . 81089:83985(2896) ack 65248 win 3125
15:01:23.546565 IP B > A: . ack 83985 win 3668
15:01:23.547713 IP A > B: . 83985:86881(2896) ack 65248 win 3125
15:01:23.547778 IP B > A: . ack 86881 win 3668
15:01:23.548911 IP A > B: . 86881:89777(2896) ack 65248 win 3125
15:01:23.548949 IP B > A: . ack 89777 win 3668
15:01:23.550116 IP A > B: . 89777:92673(2896) ack 65248 win 3125
15:01:23.550182 IP B > A: . ack 92673 win 3668
15:01:23.551333 IP A > B: . 92673:95569(2896) ack 65248 win 3125
15:01:23.551406 IP B > A: . ack 95569 win 3668
15:01:23.552539 IP A > B: . 95569:98465(2896) ack 65248 win 3125
15:01:23.552576 IP B > A: . ack 98465 win 3668
15:01:23.553756 IP A > B: . 98465:99913(1448) ack 65248 win 3125
15:01:23.554138 IP A > B: P 99913:100001(88) ack 65248 win 3125
15:01:23.554204 IP B > A: . ack 100001 win 3668
15:01:23.554234 IP B > A: . 65248:68144(2896) ack 100001 win 3668
15:01:23.555620 IP B > A: . 68144:71040(2896) ack 100001 win 3668
15:01:23.557005 IP B > A: . 71040:73936(2896) ack 100001 win 3668
15:01:23.558390 IP B > A: . 73936:76832(2896) ack 100001 win 3668
15:01:23.559773 IP B > A: . 76832:79728(2896) ack 100001 win 3668
15:01:23.561158 IP B > A: . 79728:82624(2896) ack 100001 win 3668
15:01:23.562543 IP B > A: . 82624:85520(2896) ack 100001 win 3668
15:01:23.563928 IP B > A: . 85520:88416(2896) ack 100001 win 3668
15:01:23.565313 IP B > A: . 88416:91312(2896) ack 100001 win 3668
15:01:23.566698 IP B > A: . 91312:94208(2896) ack 100001 win 3668
15:01:23.568083 IP B > A: . 94208:97104(2896) ack 100001 win 3668
15:01:23.569467 IP B > A: . 97104:100000(2896) ack 100001 win 3668
15:01:23.570852 IP B > A: . 100000:102896(2896) ack 100001 win 3668
15:01:23.572237 IP B > A: . 102896:105792(2896) ack 100001 win 3668
15:01:23.573639 IP B > A: . 105792:108688(2896) ack 100001 win 3668
15:01:23.575024 IP B > A: . 108688:111584(2896) ack 100001 win 3668
15:01:23.576408 IP B > A: . 111584:114480(2896) ack 100001 win 3668
15:01:23.577793 IP B > A: . 114480:117376(2896) ack 100001 win 3668

TCP timestamps show that most packets from B were queued in the same ms
timeframe (TSval 1159799{3,4}), but FQ managed to send them right
in time to avoid a big burst.

In slow start or steady state, very few packets are throttled [1]

FQ gets a bunch of tunables as :

limit : max number of packets on whole Qdisc (default 10000)

flow_limit : max number of packets per flow (default 100)

quantum : the credit per RR round (default is 2 MTU)

initial_quantum : initial credit for new flows (default is 10 MTU)

maxrate : max per flow rate (default : unlimited)

buckets : number of RB trees (default : 1024) in hash table.
(consumes 8 bytes per bucket)

[no]pacing : disable/enable pacing (default is enable)

All of them can be changed on a live qdisc.

$ tc qd add dev eth0 root fq help
Usage: ... fq [ limit PACKETS ] [ flow_limit PACKETS ]
[ quantum BYTES ] [ initial_quantum BYTES ]
[ maxrate RATE ] [ buckets NUMBER ]
[ [no]pacing ]

$ tc -s -d qd
qdisc fq 8002: dev eth0 root refcnt 32 limit 10000p flow_limit 100p buckets 256 quantum 3028 initial_quantum 15140
Sent 216532416 bytes 148395 pkt (dropped 0, overlimits 0 requeues 14)
backlog 0b 0p requeues 14
511 flows, 511 inactive, 0 throttled
110 gc, 0 highprio, 0 retrans, 1143 throttled, 0 flows_plimit

[1] Except if initial srtt is overestimated, as if using
cached srtt in tcp metrics. We'll provide a fix for this issue.

Signed-off-by: Eric Dumazet
Cc: Yuchung Cheng
Cc: Neal Cardwell
Signed-off-by: David S. Miller

Eric Dumazet
2013-08-30 09:38:31 +0800