Eric Lee / smarc-fsl-linux-kernel

01 Oct, 2016

1 commit

d29216842 mnt: Add a per mount namespace limit on the number of mounts ... Browse Code »

CAI Qian pointed out that the semantics
of shared subtrees make it possible to create an exponentially
increasing number of mounts in a mount namespace.

mkdir /tmp/1 /tmp/2
mount --make-rshared /
for i in $(seq 1 20) ; do mount --bind /tmp/1 /tmp/2 ; done

Will create create 2^20 or 1048576 mounts, which is a practical problem
as some people have managed to hit this by accident.

As such CVE-2016-6213 was assigned.

Ian Kent described the situation for autofs users
as follows:

> The number of mounts for direct mount maps is usually not very large because of
> the way they are implemented, large direct mount maps can have performance
> problems. There can be anywhere from a few (likely case a few hundred) to less
> than 10000, plus mounts that have been triggered and not yet expired.
>
> Indirect mounts have one autofs mount at the root plus the number of mounts that
> have been triggered and not yet expired.
>
> The number of autofs indirect map entries can range from a few to the common
> case of several thousand and in rare cases up to between 30000 and 50000. I've
> not heard of people with maps larger than 50000 entries.
>
> The larger the number of map entries the greater the possibility for a large
> number of active mounts so it's not hard to expect cases of a 1000 or somewhat
> more active mounts.

So I am setting the default number of mounts allowed per mount
namespace at 100,000. This is more than enough for any use case I
know of, but small enough to quickly stop an exponential increase
in mounts. Which should be perfect to catch misconfigurations and
malfunctioning programs.

For anyone who needs a higher limit this can be changed by writing
to the new /proc/sys/fs/mount-max sysctl.

Tested-by: CAI Qian
Signed-off-by: "Eric W. Biederman"

Eric W. Biederman
2016-10-01 01:46:48 +0800

23 Jul, 2015

1 commit

f2d0a123b mnt: Clarify and correct the disconnect logic in umount_tree ... Browse Code »

rmdir mntpoint will result in an infinite loop when there is
a mount locked on the mountpoint in another mount namespace.

This is because the logic to test to see if a mount should
be disconnected in umount_tree is buggy.

Move the logic to decide if a mount should remain connected to
it's mountpoint into it's own function disconnect_mount so that
clarity of expression instead of terseness of expression becomes
a virtue.

When the conditions where it is invalid to leave a mount connected
are first ruled out, the logic for deciding if a mount should
be disconnected becomes much clearer and simpler.

Fixes: e0c9c0afd2fc958ffa34b697972721d81df8a56f mnt: Update detach_mounts to leave mounts connected
Fixes: ce07d891a0891d3c0d0c2d73d577490486b809e1 mnt: Honor MNT_LOCKED when detaching mounts
Cc: stable@vger.kernel.org
Signed-off-by: "Eric W. Biederman"

Eric W. Biederman
2015-07-23 09:33:27 +0800

10 Apr, 2015

1 commit

ce07d891a mnt: Honor MNT_LOCKED when detaching mounts ... Browse Code »

Modify umount(MNT_DETACH) to keep mounts in the hash table that are
locked to their parent mounts, when the parent is lazily unmounted.

In mntput_no_expire detach the children from the hash table, depending
on mnt_pin_kill in cleanup_mnt to decrement the mnt_count of the children.

In __detach_mounts if there are any mounts that have been unmounted
but still are on the list of mounts of a mountpoint, remove their
children from the mount hash table and those children to the unmounted
list so they won't linger potentially indefinitely waiting for their
final mntput, now that the mounts serve no purpose.

Cc: stable@vger.kernel.org
Signed-off-by: "Eric W. Biederman"

Eric W. Biederman
2015-04-10 00:39:55 +0800

03 Apr, 2015

4 commits

0c56fe314 mnt: Don't propagate unmounts to locked mounts ... Browse Code »

If the first mount in shared subtree is locked don't unmount the
shared subtree.

This is ensured by walking through the mounts parents before children
and marking a mount as unmountable if it is not locked or it is locked
but it's parent is marked.

This allows recursive mount detach to propagate through a set of
mounts when unmounting them would not reveal what is under any locked
mount.

Cc: stable@vger.kernel.org
Signed-off-by: "Eric W. Biederman"

Eric W. Biederman
2015-04-03 09:34:20 +0800
5d88457eb mnt: On an unmount propagate clearing of MNT_LOCKED ... Browse Code »

A prerequisite of calling umount_tree is that the point where the tree
is mounted at is valid to unmount.

If we are propagating the effect of the unmount clear MNT_LOCKED in
every instance where the same filesystem is mounted on the same
mountpoint in the mount tree, as we know (by virtue of the fact
that umount_tree was called) that it is safe to reveal what
is at that mountpoint.

Cc: stable@vger.kernel.org
Signed-off-by: "Eric W. Biederman"

Eric W. Biederman
2015-04-03 09:34:19 +0800
c003b26ff mnt: In umount_tree reuse mnt_list instead of mnt_hash ... Browse Code »

umount_tree builds a list of mounts that need to be unmounted.
Utilize mnt_list for this purpose instead of mnt_hash. This begins to
allow keeping a mount on the mnt_hash after it is unmounted, which is
necessary for a properly functioning MNT_LOCKED implementation.

The fact that mnt_list is an ordinary list makding available list_move
is nice bonus.

Cc: stable@vger.kernel.org
Signed-off-by: "Eric W. Biederman"

Eric W. Biederman
2015-04-03 09:34:18 +0800
e819f1521 mnt: Improve the umount_tree flags ... Browse Code »

- Remove the unneeded declaration from pnode.h
- Mark umount_tree static as it has no callers outside of namespace.c
- Define an enumeration of umount_tree's flags.
- Pass umount_tree's flags in by name

This removes the magic numbers 0, 1 and 2 making the code a little
clearer and makes it possible for there to be lazy unmounts that don't
propagate. Which is what __detach_mounts actually wants for example.

Cc: stable@vger.kernel.org
Signed-off-by: "Eric W. Biederman"

Eric W. Biederman
2015-04-03 09:34:17 +0800

02 Apr, 2014

1 commit

f2ebb3a92 smarter propagate_mnt() ... Browse Code »

The current mainline has copies propagated to *all* nodes, then
tears down the copies we made for nodes that do not contain
counterparts of the desired mountpoint. That sets the right
propagation graph for the copies (at teardown time we move
the slaves of removed node to a surviving peer or directly
to master), but we end up paying a fairly steep price in
useless allocations. It's fairly easy to create a situation
where N calls of mount(2) create exactly N bindings, with
O(N^2) vfsmounts allocated and freed in process.

Fortunately, it is possible to avoid those allocations/freeings.
The trick is to create copies in the right order and find which
one would've eventually become a master with the current algorithm.
It turns out to be possible in O(nodes getting propagation) time
and with no extra allocations at all.

One part is that we need to make sure that eventual master will be
created before its slaves, so we need to walk the propagation
tree in a different order - by peer groups. And iterate through
the peers before dealing with the next group.

Another thing is finding the (earlier) copy that will be a master
of one we are about to create; to do that we are (temporary) marking
the masters of mountpoints we are attaching the copies to.

Either we are in a peer of the last mountpoint we'd dealt with,
or we have the following situation: we are attaching to mountpoint M,
the last copy S_0 had been attached to M_0 and there are sequences
S_0...S_n, M_0...M_n such that S_{i+1} is a master of S_{i},
S_{i} mounted on M{i} and we need to create a slave of the first S_{k}
such that M is getting propagation from M_{k}. It means that the master
of M_{k} will be among the sequence of masters of M. On the
other hand, the nearest marked node in that sequence will either
be the master of M_{k} or the master of M_{k-1} (the latter -
in the case if M_{k-1} is a slave of something M gets propagation
from, but in a wrong peer group).

So we go through the sequence of masters of M until we find
a marked one (P). Let N be the one before it. Then we go through
the sequence of masters of S_0 until we find one (say, S) mounted
on a node D that has P as master and check if D is a peer of N.
If it is, S will be the master of new copy, if not - the master of S
will be.

That's it for the hard part; the rest is fairly simple. Iterator
is in next_group(), handling of one prospective mountpoint is
propagate_one().

It seems to survive all tests and gives a noticably better performance
than the current mainline for setups that are seriously using shared
subtrees.

Cc: stable@vger.kernel.org
Signed-off-by: Al Viro

Al Viro
2014-04-02 11:19:08 +0800

31 Mar, 2014

1 commit

38129a13e switch mnt_hash to hlist ... Browse Code »

fixes RCU bug - walking through hlist is safe in face of element moves,
since it's self-terminating. Cyclic lists are not - if we end up jumping
to another hash chain, we'll loop infinitely without ever hitting the
original list head.

[fix for dumb braino folded]

Spotted by: Max Kellermann
Cc: stable@vger.kernel.org
Signed-off-by: Al Viro

Al Viro
2014-03-31 07:18:51 +0800

27 Aug, 2013

1 commit

4ce5d2b1a vfs: Don't copy mount bind mounts of /proc/<pid>/ns/mnt between namespaces ... Browse Code »

Don't copy bind mounts of /proc//ns/mnt between namespaces.
These files hold references to a mount namespace and copying them
between namespaces could result in a reference counting loop.

The current mnt_ns_loop test prevents loops on the assumption that
mounts don't cross between namespaces. Unfortunately unsharing a
mount namespace and shared substrees can both cause mounts to
propogate between mount namespaces.

Add two flags CL_COPY_UNBINDABLE and CL_COPY_MNT_NS_FILE are added to
control this behavior, and CL_COPY_ALL is redefined as both of them.

Signed-off-by: "Eric W. Biederman"

Eric W. Biederman
2013-08-27 09:42:15 +0800

02 May, 2013

1 commit

20b4fb485 Merge branch 'for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/viro/vfs ... Browse Code »

Pull VFS updates from Al Viro,

Misc cleanups all over the place, mainly wrt /proc interfaces (switch
create_proc_entry to proc_create(), get rid of the deprecated
create_proc_read_entry() in favor of using proc_create_data() and
seq_file etc).

7kloc removed.

* 'for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/viro/vfs: (204 commits)
don't bother with deferred freeing of fdtables
proc: Move non-public stuff from linux/proc_fs.h to fs/proc/internal.h
proc: Make the PROC_I() and PDE() macros internal to procfs
proc: Supply a function to remove a proc entry by PDE
take cgroup_open() and cpuset_open() to fs/proc/base.c
ppc: Clean up scanlog
ppc: Clean up rtas_flash driver somewhat
hostap: proc: Use remove_proc_subtree()
drm: proc: Use remove_proc_subtree()
drm: proc: Use minor->index to label things, not PDE->name
drm: Constify drm_proc_list[]
zoran: Don't print proc_dir_entry data in debug
reiserfs: Don't access the proc_dir_entry in r_open(), r_start() r_show()
proc: Supply an accessor for getting the data from a PDE's parent
airo: Use remove_proc_subtree()
rtl8192u: Don't need to save device proc dir PDE
rtl8187se: Use a dir under /proc/net/r8180/
proc: Add proc_mkdir_data()
proc: Move some bits from linux/proc_fs.h to linux/{of.h,signal.h,tty.h}
proc: Move PDE_NET() to fs/proc/proc_net.c
...

Linus Torvalds
2013-05-02 08:51:54 +0800

10 Apr, 2013

2 commits

328e6d901 switch unlock_mount() to namespace_unlock(), convert all umount_tree() callers ... Browse Code »

which allows to kill the last argument of umount_tree() and make release_mounts()
static.

Signed-off-by: Al Viro

Al Viro
2013-04-10 02:12:53 +0800
84d17192d get rid of full-hash scan on detaching vfsmounts ... Browse Code »

Signed-off-by: Al Viro

Al Viro
2013-04-10 02:12:52 +0800

27 Mar, 2013

1 commit

132c94e31 vfs: Carefully propogate mounts across user namespaces ... Browse Code »

As a matter of policy MNT_READONLY should not be changable if the
original mounter had more privileges than creator of the mount
namespace.

Add the flag CL_UNPRIVILEGED to note when we are copying a mount from
a mount namespace that requires more privileges to a mount namespace
that requires fewer privileges.

When the CL_UNPRIVILEGED flag is set cause clone_mnt to set MNT_NO_REMOUNT
if any of the mnt flags that should never be changed are set.

This protects both mount propagation and the initial creation of a less
privileged mount namespace.

Cc: stable@vger.kernel.org
Acked-by: Serge Hallyn
Reported-by: Andy Lutomirski
Signed-off-by: "Eric W. Biederman"

Eric W. Biederman
2013-03-27 22:50:05 +0800

19 Nov, 2012

1 commit

7a472ef4b vfs: Only support slave subtrees across different user namespaces ... Browse Code »

Sharing mount subtress with mount namespaces created by unprivileged
users allows unprivileged mounts created by unprivileged users to
propagate to mount namespaces controlled by privileged users.

Prevent nasty consequences by changing shared subtrees to slave
subtress when an unprivileged users creates a new mount namespace.

Acked-by: Serge Hallyn
Signed-off-by: "Eric W. Biederman"

Eric W. Biederman
2012-11-19 21:59:20 +0800

04 Jan, 2012

17 commits

fc7be130c vfs: switch pnode.h macros to struct mount * ... Browse Code »

Signed-off-by: Al Viro

Al Viro
2012-01-04 11:57:11 +0800
14cf1fa8f vfs: spread struct mount - remaining argument of mnt_set_mountpoint() ... Browse Code »

Signed-off-by: Al Viro

Al Viro
2012-01-04 11:57:07 +0800
a8d56d8e4 vfs: spread struct mount - propagate_mnt() ... Browse Code »

Signed-off-by: Al Viro

Al Viro
2012-01-04 11:57:07 +0800
6fc7871fe vfs: spread struct mount - get_dominating_id / do_make_slave ... Browse Code »

next pile of horrors, similar to mnt_parent one; this time it's
mnt_master.

Signed-off-by: Al Viro

Al Viro
2012-01-04 11:57:06 +0800
83adc7532 vfs: spread struct mount - work with counters ... Browse Code »

Signed-off-by: Al Viro

Al Viro
2012-01-04 11:57:05 +0800
643822b41 vfs: spread struct mount - is_path_reachable ... Browse Code »

Signed-off-by: Al Viro

Al Viro
2012-01-04 11:57:04 +0800
1ab597386 vfs: spread struct mount - do_umount/propagate_mount_busy ... Browse Code »

Signed-off-by: Al Viro

Al Viro
2012-01-04 11:57:03 +0800
44d964d60 vfs: spread struct mount mnt_set_mountpoint child argument ... Browse Code »

Signed-off-by: Al Viro

Al Viro
2012-01-04 11:57:03 +0800
87129cc0e vfs: spread struct mount - clone_mnt/copy_tree argument ... Browse Code »

Signed-off-by: Al Viro

Al Viro
2012-01-04 11:57:03 +0800
761d5c38e vfs: spread struct mount - umount_tree argument ... Browse Code »

Signed-off-by: Al Viro

Al Viro
2012-01-04 11:57:02 +0800
cb338d06e vfs: spread struct mount - clone_mnt/copy_tree result ... Browse Code »

Signed-off-by: Al Viro

Al Viro
2012-01-04 11:57:01 +0800
0f0afb1dc vfs: spread struct mount - change_mnt_propagation/set_mnt_shared ... Browse Code »

Signed-off-by: Al Viro

Al Viro
2012-01-04 11:57:01 +0800
4b8b21f4f vfs: spread struct mount - mount group id handling ... Browse Code »

Signed-off-by: Al Viro

Al Viro
2012-01-04 11:56:59 +0800
aafd08dad vfs: add missing parens in pnode.h macros ... Browse Code »

Signed-off-by: Al Viro

Al Viro
2012-01-04 11:52:37 +0800
afac7cba7 vfs: more mnt_parent cleanups ... Browse Code »

a) mount --move is checking that ->mnt_parent is non-NULL before
looking if that parent happens to be shared; ->mnt_parent is never
NULL and it's not even an misspelled !mnt_has_parent()

b) pivot_root open-codes is_path_reachable(), poorly.

c) so does path_is_under(), while we are at it.

Signed-off-by: Al Viro

Al Viro
2012-01-04 11:52:36 +0800
b2dba1af3 vfs: new internal helper: mnt_has_parent(mnt) ... Browse Code »

vfsmounts have ->mnt_parent pointing either to a different vfsmount
or to itself; it's never NULL and termination condition in loops
traversing the tree towards root is mnt == mnt->mnt_parent. At least
one place (see the next patch) is confused about what's going on;
let's add an explicit helper checking it right way and use it in
all places where we need it. Not that there had been too many,
but...

Signed-off-by: Al Viro

Al Viro
2012-01-04 11:52:36 +0800
f47ec3f28 trim fs/internal.h ... Browse Code »

some stuff in there can actually become static; some belongs to pnode.h
as it's a private interface between namespace.c and pnode.c...

Signed-off-by: Al Viro

Al Viro
2012-01-04 11:52:35 +0800

04 Mar, 2010

2 commits

495d6c9c6 VFS: Clean up shared mount flag propagation ... Browse Code »

The handling of mount flags in set_mnt_shared() got a little tangled
up during previous cleanups, with the following problems:

* MNT_PNODE_MASK is defined as a literal constant when it should be a
bitwise xor of other MNT_* flags
* set_mnt_shared() clears and then sets MNT_SHARED (part of MNT_PNODE_MASK)
* MNT_PNODE_MASK could use a comment in mount.h
* MNT_PNODE_MASK is a terrible name, change to MNT_SHARED_MASK

This patch fixes these problems.

Signed-off-by: Al Viro

Valerie Aurora
2010-03-04 03:07:55 +0800
796a6b521 Kill CL_PROPAGATION, sanitize fs/pnode.c:get_source() ... Browse Code »

First of all, get_source() never results in CL_PROPAGATION
alone. We either get CL_MAKE_SHARED (for the continuation
of peer group) or CL_SLAVE (slave that is not shared) or both
(beginning of peer group among slaves). Massage the code to
make that explicit, kill CL_PROPAGATION test in clone_mnt()
(nothing sets CL_MAKE_SHARED without CL_PROPAGATION and in
clone_mnt() we are checking CL_PROPAGATION after we'd found
that there's no CL_SLAVE, so the check for CL_MAKE_SHARED
would do just as well).

Fix comments, while we are at it...

Signed-off-by: Al Viro

Al Viro
2010-03-04 02:00:22 +0800

23 Apr, 2008

1 commit

97e7e0f71 [patch 7/7] vfs: mountinfo: show dominating group id ... Browse Code »

Show peer group ID of nearest dominating group that has intersection
with the mount's namespace.

Signed-off-by: Miklos Szeredi
Signed-off-by: Al Viro

Miklos Szeredi
2008-04-23 12:05:09 +0800

22 Apr, 2008

1 commit

521b5d0c4 [PATCH] teach seq_file to discard entries ... Browse Code »

Allow ->show() return SEQ_SKIP; that will discard all
output from that element and move on.

Signed-off-by: Al Viro

Al Viro
2008-04-22 11:14:02 +0800

21 Oct, 2007

1 commit

8aec08094 [PATCH] new helpers - collect_mounts() and release_collected_mounts() ... Browse Code »

Get a snapshot of a subtree, creating private clones of vfsmounts
for all its components and release such snapshot resp.

Signed-off-by: Al Viro

Al Viro
2007-10-21 14:37:25 +0800

09 Dec, 2006

1 commit

6b3286ed1 [PATCH] rename struct namespace to struct mnt_namespace ... Browse Code »

Rename 'struct namespace' to 'struct mnt_namespace' to avoid confusion with
other namespaces being developped for the containers : pid, uts, ipc, etc.
'namespace' variables and attributes are also renamed to 'mnt_ns'

Signed-off-by: Kirill Korotaev
Signed-off-by: Cedric Le Goater
Cc: Eric W. Biederman
Cc: Herbert Poetzl
Cc: Sukadev Bhattiprolu
Signed-off-by: Andrew Morton
Signed-off-by: Linus Torvalds

Kirill Korotaev
2006-12-09 00:28:51 +0800

08 Nov, 2005

2 commits

9676f0c63 [PATCH] unbindable mounts ... Browse Code »

An unbindable mount does not forward or receive propagation. Also
unbindable mount disallows bind mounts. The semantics is as follows.

Bind semantics:
It is invalid to bind mount an unbindable mount.

Move semantics:
It is invalid to move an unbindable mount under shared mount.

Clone-namespace semantics:
If a mount is unbindable in the parent namespace, the corresponding
cloned mount in the child namespace becomes unbindable too. Note:
there is subtle difference, unbindable mounts cannot be bind mounted
but can be cloned during clone-namespace.

Signed-off-by: Ram Pai
Signed-off-by: Al Viro
Signed-off-by: Linus Torvalds

Ram Pai
2005-11-08 10:18:11 +0800
a58b0eb8e [PATCH] introduce slave mounts ... Browse Code »

A slave mount always has a master mount from which it receives
mount/umount events. Unlike shared mount the event propagation does not
flow from the slave mount to the master.

Signed-off-by: Ram Pai
Signed-off-by: Al Viro
Signed-off-by: Linus Torvalds

Ram Pai
2005-11-08 10:18:11 +0800