Eric Lee / smarc-fsl-linux-kernel

05 Jul, 2013

1 commit

80cc38b16 Merge branch 'for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/jikos/trivial ... Browse Code »

Pull trivial tree updates from Jiri Kosina:
"The usual stuff from trivial tree"

* 'for-linus' of git://git.kernel.org/pub/scm/linux/kernel/git/jikos/trivial: (34 commits)
treewide: relase -> release
Documentation/cgroups/memory.txt: fix stat file documentation
sysctl/net.txt: delete reference to obsolete 2.4.x kernel
spinlock_api_smp.h: fix preprocessor comments
treewide: Fix typo in printk
doc: device tree: clarify stuff in usage-model.txt.
open firmware: "/aliasas" -> "/aliases"
md: bcache: Fixed a typo with the word 'arithmetic'
irq/generic-chip: fix a few kernel-doc entries
frv: Convert use of typedef ctl_table to struct ctl_table
sgi: xpc: Convert use of typedef ctl_table to struct ctl_table
doc: clk: Fix incorrect wording
Documentation/arm/IXP4xx fix a typo
Documentation/networking/ieee802154 fix a typo
Documentation/DocBook/media/v4l fix a typo
Documentation/video4linux/si476x.txt fix a typo
Documentation/virtual/kvm/api.txt fix a typo
Documentation/early-userspace/README fix a typo
Documentation/video4linux/soc-camera.txt fix a typo
lguest: fix CONFIG_PAE -> CONFIG_x86_PAE in comment
...

Linus Torvalds
2013-07-05 02:40:58 +0800

04 Jul, 2013

2 commits

137651206 md/raid10: fix bug which causes all RAID10 reshapes to move no data. ... Browse Code »

The recent comment:
commit 7e83ccbecd608b971f340e951c9e84cd0343002f
md/raid10: Allow skipping recovery when clean arrays are assembled

Causes raid10 to skip a recovery in certain cases where it is safe to
do so. Unfortunately it also causes a reshape to be skipped which is
never safe. The result is that an attempt to reshape a RAID10 will
appear to complete instantly, but no data will have been moves so the
array will now contain garbage.
(If nothing is written, you can recovery by simple performing the
reverse reshape which will also complete instantly).

Bug was introduced in 3.10, so this is suitable for 3.10-stable.

Cc: stable@vger.kernel.org (3.10)
Cc: Martin Wilck
Signed-off-by: NeilBrown

NeilBrown
2013-07-04 14:42:57 +0800
fdcfbbb65 md/raid5: allow 5-device RAID6 to be reshaped to 4-device. ... Browse Code »

There is a bug in 'check_reshape' for raid5.c To checks
that the new minimum number of devices is large enough (which is
good), but it does so also after the reshape has started (bad).

This is bad because
- the calculation is now wrong as mddev->raid_disks has changed
already, and
- it is pointless because it is now too late to stop.

So only perform that test when reshape has not been committed to.

Signed-off-by: NeilBrown

NeilBrown
2013-07-04 14:42:52 +0800

03 Jul, 2013

1 commit

78eaa0d4c md/raid10: fix two bugs affecting RAID10 reshape. ... Browse Code »

1/ If a RAID10 is being reshaped to a fewer number of devices
and is stopped while this is ongoing, then when the array is
reassembled the 'mirrors' array will be allocated too small.
This will lead to an access error or memory corruption.

2/ A sanity test for a reshaping RAID10 array is restarted
is slightly incorrect.

Due to the first bug, this is suitable for any -stable
kernel since 3.5 where this code was introduced.

Cc: stable@vger.kernel.org (v3.5+)
Signed-off-by: NeilBrown

NeilBrown
2013-07-03 07:43:28 +0800

26 Jun, 2013

2 commits

c4a395514 MD: Remember the last sync operation that was performed ... Browse Code »

MD: Remember the last sync operation that was performed

This patch adds a field to the mddev structure to track the last
sync operation that was performed. This is especially useful when
it comes to what is recorded in mismatch_cnt in sysfs. If the
last operation was "data-check", then it reports the number of
descrepancies found by the user-initiated check. If it was a
"repair" operation, then it is reporting the number of
descrepancies repaired. etc.

Signed-off-by: Jonathan Brassow
Signed-off-by: NeilBrown

Jonathan Brassow
2013-06-26 10:38:24 +0800
eea136d69 md: fix buglet in RAID5 -> RAID0 conversion. ... Browse Code »

RAID5 uses a 'per-array' value for the 'size' of each device.
RAID0 uses a 'per-device' value - it can be different for each device.

When converting a RAID5 to a RAID0 we must ensure that the per-device
size of each device matches the per-array size for the RAID5, else
the array will change size.

If the metadata cannot record a changed per-device size (as is the
case with v0.90 metadata) the array could get bigger on restart. This
does not cause data corruption, so it not a big issue and is mainly
yet another a reason to not use 0.90.

Signed-off-by: NeilBrown

NeilBrown
2013-06-26 10:38:19 +0800

18 Jun, 2013

1 commit

48a73025c md: bcache: Fixed a typo with the word 'arithmetic' ... Browse Code »

The word 'arithmetic' was typed as 'arithmatic'

Signed-off-by: Phil Viana
Signed-off-by: Jiri Kosina

Phil Viana
2013-06-18 19:41:16 +0800

14 Jun, 2013

9 commits

725d6e579 md/raid10: check In_sync flag in 'enough()'. ... Browse Code »

It isn't really enough to check that the rdev is present, we need to
also be sure that the device is still In_sync.

Doing this requires using rcu_dereference to access the rdev, and
holding the rcu_read_lock() to ensure the rdev doesn't disappear while
we look at it.

Signed-off-by: NeilBrown

NeilBrown
2013-06-14 06:10:27 +0800
635f6416a md/raid10: locking changes for 'enough()'. ... Browse Code »

As 'enough' accesses conf->prev and conf->geo, which can change
spontanously, it should guard against changes.
This can be done with device_lock as start_reshape holds device_lock
while updating 'geo' and end_reshape holds it while updating 'prev'.

So 'error' needs to hold 'device_lock'.

On the other hand, raid10_end_read_request knows which of the two it
really wants to access, and as it is an active request on that one,
the value cannot change underneath it.

So change _enough to take flag rather than a pointer, pass the
appropriate flag from raid10_end_read_request(), and remove the locking.

All other calls to 'enough' are made with reconfig_mutex held, so
neither 'prev' nor 'geo' can change.

Signed-off-by: NeilBrown

NeilBrown
2013-06-14 06:10:27 +0800
b29bebd66 md: replace strict_strto*() with kstrto*() ... Browse Code »

The usage of strict_strtoul() is not preferred, because
strict_strtoul() is obsolete. Thus, kstrtoul() should be
used.

Signed-off-by: Jingoo Han
Signed-off-by: NeilBrown

Jingoo Han
2013-06-14 06:10:26 +0800
90f5f7ad4 md: Wait for md_check_recovery before attempting device removal. ... Browse Code »

When a device has failed, it needs to be removed from the personality
module before it can be removed from the array as a whole.
The first step is performed by md_check_recovery() which is called
from the raid management thread.

So when a HOT_REMOVE ioctl arrives, wait briefly for md_check_recovery
to have run. This increases the chance that the ioctl will succeed.

Signed-off-by: Hannes Reinecke
Signed-off-by: Neil Brown

Hannes Reinecke
2013-06-14 06:10:26 +0800
3f6bbd3ff dm-raid: silence compiler warning on rebuilds_per_group. ... Browse Code »

This doesn't really need to be initialised, but it doesn't hurt,
silences the compiler, and as it is a counter it makes sense for it to
start at zero.

Signed-off-by: NeilBrown

NeilBrown
2013-06-14 06:10:26 +0800
a4dc163a5 DM RAID: Fix raid_resume not reviving failed devices in all cases ... Browse Code »

DM RAID: Fix raid_resume not reviving failed devices in all cases

When a device fails in a RAID array, it is marked as Faulty. Later,
md_check_recovery is called which (through the call chain) calls
'hot_remove_disk' in order to have the personalities remove the device
from use in the array.

Sometimes, it is possible for the array to be suspended before the
personalities get their chance to perform 'hot_remove_disk'. This is
normally not an issue. If the array is deactivated, then the failed
device will be noticed when the array is reinstantiated. If the
array is resumed and the disk is still missing, md_check_recovery will
be called upon resume and 'hot_remove_disk' will be called at that
time. However, (for dm-raid) if the device has been restored,
a resume on the array would cause it to attempt to revive the device
by calling 'hot_add_disk'. If 'hot_remove_disk' had not been called,
a situation is then created where the device is thought to concurrently
be the replacement and the device to be replaced. Thus, the device
is first sync'ed with the rest of the array (because it is the replacement
device) and then marked Faulty and removed from the array (because
it is also the device being replaced).

The solution is to check and see if the device had properly been removed
before the array was suspended. This is done by seeing whether the
device's 'raid_disk' field is -1 - a condition that implies that
'md_check_recovery -> remove_and_add_spares (where raid_disk is set to -1)
-> hot_remove_disk' has been called. If 'raid_disk' is not -1, then
'hot_remove_disk' must be called to complete the removal of the previously
faulty device before it can be revived via 'hot_add_disk'.

Signed-off-by: Jonathan Brassow
Signed-off-by: NeilBrown

Jonathan Brassow
2013-06-14 06:10:25 +0800
f381e71b0 DM RAID: Break-up untidy function ... Browse Code »

DM RAID: Break-up untidy function

Clean-up excessive indentation by moving some code in raid_resume()
into its own function.

Signed-off-by: Jonathan Brassow
Signed-off-by: NeilBrown

Jonathan Brassow
2013-06-14 06:10:25 +0800
9092c02d9 DM RAID: Add ability to restore transiently failed devices on resume ... Browse Code »

DM RAID: Add ability to restore transiently failed devices on resume

This patch adds code to the resume function to check over the devices
in the RAID array. If any are found to be marked as failed and their
superblocks can be read, an attempt is made to reintegrate them into
the array. This allows the user to refresh the array with a simple
suspend and resume of the array - rather than having to load a
completely new table, allocate and initialize all the structures and
throw away the old instantiation.

Signed-off-by: Jonathan Brassow
Signed-off-by: NeilBrown

Jonathan Brassow
2013-06-14 06:10:24 +0800
82ea4be61 Merge tag 'md-3.10-fixes' of git://neil.brown.name/md ... Browse Code »

Pull md bugfixes from Neil Brown:
"A few bugfixes for md

Some tagged for -stable"

* tag 'md-3.10-fixes' of git://neil.brown.name/md:
md/raid1,5,10: Disable WRITE SAME until a recovery strategy is in place
md/raid1,raid10: use freeze_array in place of raise_barrier in various places.
md/raid1: consider WRITE as successful only if at least one non-Faulty and non-rebuilding drive completed it.
md: md_stop_writes() should always freeze recovery.

Linus Torvalds
2013-06-14 01:13:29 +0800

13 Jun, 2013

5 commits

5026d7a9b md/raid1,5,10: Disable WRITE SAME until a recovery strategy is in place ... Browse Code »

There are cases where the kernel will believe that the WRITE SAME
command is supported by a block device which does not, in fact,
support WRITE SAME. This currently happens for SATA drivers behind a
SAS controller, but there are probably a hundred other ways that can
happen, including drive firmware bugs.

After receiving an error for WRITE SAME the block layer will retry the
request as a plain write of zeroes, but mdraid will consider the
failure as fatal and consider the drive failed. This has the effect
that all the mirrors containing a specific set of data are each
offlined in very rapid succession resulting in data loss.

However, just bouncing the request back up to the block layer isn't
ideal either, because the whole initial request-retry sequence should
be inside the write bitmap fence, which probably means that md needs
to do its own conversion of WRITE SAME to write zero.

Until the failure scenario has been sorted out, disable WRITE SAME for
raid1, raid5, and raid10.

[neilb: added raid5]

This patch is appropriate for any -stable since 3.7 when write_same
support was added.

Cc: stable@vger.kernel.org
Signed-off-by: H. Peter Anvin
Signed-off-by: NeilBrown

H. Peter Anvin
2013-06-13 12:49:54 +0800
e2d599252 md/raid1,raid10: use freeze_array in place of raise_barrier in various places. ... Browse Code »

Various places in raid1 and raid10 are calling raise_barrier when they
really should call freeze_array.
The former is only intended to be called from "make_request".
The later has extra checks for 'nr_queued' and makes a call to
flush_pending_writes(), so it is safe to call it from within the
management thread.

Using raise_barrier will sometimes deadlock. Using freeze_array
should not.

As 'freeze_array' currently expects one request to be pending (in
handle_read_error - the only previous caller), we need to pass
it the number of pending requests (extra) to ignore.

The deadlock was made particularly noticeable by commits
050b66152f87c7 (raid10) and 6b740b8d79252f13 (raid1) which
appeared in 3.4, so the fix is appropriate for any -stable
kernel since then.

This patch probably won't apply directly to some early kernels and
will need to be applied by hand.

Cc: stable@vger.kernel.org
Reported-by: Alexander Lyakas
Signed-off-by: NeilBrown

NeilBrown
2013-06-13 11:40:48 +0800
3056e3aec md/raid1: consider WRITE as successful only if at least one non-Faulty and non-r… ... Browse Code »

…ebuilding drive completed it.

Without that fix, the following scenario could happen:

- RAID1 with drives A and B; drive B was freshly-added and is rebuilding
- Drive A fails
- WRITE request arrives to the array. It is failed by drive A, so
r1_bio is marked as R1BIO_WriteError, but the rebuilding drive B
succeeds in writing it, so the same r1_bio is marked as
R1BIO_Uptodate.
- r1_bio arrives to handle_write_finished, badblocks are disabled,
md_error()->error() does nothing because we don't fail the last drive
of raid1
- raid_end_bio_io() calls call_bio_endio()
- As a result, in call_bio_endio():
if (!test_bit(R1BIO_Uptodate, &r1_bio->state))
clear_bit(BIO_UPTODATE, &bio->bi_flags);
this code doesn't clear the BIO_UPTODATE flag, and the whole master
WRITE succeeds, back to the upper layer.

So we returned success to the upper layer, even though we had written
the data onto the rebuilding drive only. But when we want to read the
data back, we would not read from the rebuilding drive, so this data
is lost.

[neilb - applied identical change to raid10 as well]

This bug can result in lost data, so it is suitable for any
-stable kernel.

Cc: stable@vger.kernel.org
Signed-off-by: Alex Lyakas <alex@zadarastorage.com>
Signed-off-by: NeilBrown <neilb@suse.de>

Alex Lyakas
2013-06-13 11:20:03 +0800
6b6204ee9 md: md_stop_writes() should always freeze recovery. ... Browse Code »

__md_stop_writes() will currently sometimes freeze recovery.
So any caller must be ready for that to happen, and indeed they are.

However if __md_stop_writes() doesn't freeze_recovery, then
a recovery could start before mddev_suspend() is called, which
could be awkward. This can particularly cause problems or dm-raid.

So change __md_stop_writes() to always freeze recovery. This is safe
and more predicatable.

Reported-by: Brassow Jonathan
Tested-by: Brassow Jonathan
Signed-off-by: NeilBrown

NeilBrown
2013-06-13 11:18:15 +0800
b2cc9c19e Merge branch 'for-linus' of git://git.kernel.dk/linux-block ... Browse Code »

Pull block layer fixes from Jens Axboe:
"Outside of bcache (which really isn't super big), these are all
few-liners. There are a few important fixes in here:

- Fix blk pm sleeping when holding the queue lock

- A small collection of bcache fixes that have been done and tested
since bcache was included in this merge window.

- A fix for a raid5 regression introduced with the bio changes.

- Two important fixes for mtip32xx, fixing an oops and potential data
corruption (or hang) due to wrong bio iteration on stacked devices."

* 'for-linus' of git://git.kernel.dk/linux-block:
scatterlist: sg_set_buf() argument must be in linear mapping
raid5: Initialize bi_vcnt
pktcdvd: silence static checker warning
block: remove refs to XD disks from documentation
blkpm: avoid sleep when holding queue lock
mtip32xx: Correctly handle bio->bi_idx != 0 conditions
mtip32xx: Fix NULL pointer dereference during module unload
bcache: Fix error handling in init code
bcache: clarify free/available/unused space
bcache: drop "select CLOSURES"
bcache: Fix incompatible pointer type warning

Linus Torvalds
2013-06-13 07:42:39 +0800

30 May, 2013

1 commit

4997b72ee raid5: Initialize bi_vcnt ... Browse Code »

The patch that converted raid5 to use bio_reset() forgot to initialize
bi_vcnt.

Signed-off-by: Kent Overstreet
Cc: NeilBrown
Cc: linux-raid@vger.kernel.org
Tested-by: Ilia Mirkin
Signed-off-by: Jens Axboe

Kent Overstreet
2013-05-30 14:44:39 +0800

20 May, 2013

1 commit

610bba8b9 dm thin: fix metadata dev resize detection ... Browse Code »

Fix detection of the need to resize the dm thin metadata device.

The code incorrectly tried to extend the metadata device when it
didn't need to due to a merging error with patch 24347e9 ("dm thin:
detect metadata device resizing").

device-mapper: transaction manager: couldn't open metadata space map
device-mapper: thin metadata: tm_open_with_sm failed
device-mapper: thin: aborting transaction failed
device-mapper: thin: switching pool to failure mode

Signed-off-by: Alasdair G Kergon

Alasdair G Kergon
2013-05-20 01:57:50 +0800

15 May, 2013

4 commits

c0a363f5c Merge branch 'bcache-for-upstream' of git://evilpiepirate.org/~kent/linux-bcache into for-linus ... Browse Code »

Kent writes:

Jens - couple more bcache patches. Bug fixes and a doc update.

Jens Axboe
2013-05-15 16:36:25 +0800
f59fce847 bcache: Fix error handling in init code ... Browse Code »

This code appears to have rotted... fix various bugs and do some
refactoring.

Signed-off-by: Kent Overstreet

Kent Overstreet
2013-05-15 15:48:14 +0800
bbb1c3b5a bcache: drop "select CLOSURES" ... Browse Code »

The Kconfig entry for BCACHE selects CLOSURES. But there's no Kconfig
symbol CLOSURES. That symbol was used in development versions of bcache,
but was removed when the closures code was no longer provided as a
kernel library. It can safely be dropped.

Signed-off-by: Paul Bolle

Paul Bolle
2013-05-15 15:42:51 +0800
867e11620 bcache: Fix incompatible pointer type warning ... Browse Code »

The function pointer release in struct block_device_operations
should point to functions declared as void.

Sparse warnings:

drivers/md/bcache/super.c:656:27: warning:
incorrect type in initializer (different base types)
drivers/md/bcache/super.c:656:27:
expected void ( *release )( ... )
drivers/md/bcache/super.c:656:27:
got int ( static [toplevel] * )( ... )

drivers/md/bcache/super.c:656:2: warning:
initialization from incompatible pointer type [enabled by default]

drivers/md/bcache/super.c:656:2: warning:
(near initialization for ‘bcache_ops.release’) [enabled by default]

Signed-off-by: Emil Goode
Signed-off-by: Kent Overstreet

Emil Goode
2013-05-15 15:42:50 +0800

10 May, 2013

13 commits

2f14f4b51 dm cache: set config value ... Browse Code »

Share configuration option processing code between the dm cache
ctr and message functions.

Signed-off-by: Joe Thornber
Signed-off-by: Alasdair G Kergon

Joe Thornber
2013-05-10 21:37:21 +0800
2c73c471f dm cache: move config fns ... Browse Code »

Move process_config_option() in dm-cache-target.c to make the
next patch more readable.

Signed-off-by: Alasdair G Kergon

Alasdair G Kergon
2013-05-10 21:37:21 +0800
ac8c3f3df dm thin: generate event when metadata threshold passed ... Browse Code »

Generate a dm event when the amount of remaining thin pool metadata
space falls below a certain level.

The threshold is taken to be a quarter of the size of the metadata
device with a minimum threshold of 4MB.

Signed-off-by: Joe Thornber
Signed-off-by: Alasdair G Kergon

Joe Thornber
2013-05-10 21:37:21 +0800
2fc48021f dm persistent metadata: add space map threshold callback ... Browse Code »

Add a threshold callback to dm persistent data space maps.

Signed-off-by: Joe Thornber
Signed-off-by: Alasdair G Kergon

Joe Thornber
2013-05-10 21:37:20 +0800
7c3d3f2a8 dm persistent data: add threshold callback to space map ... Browse Code »

Add a threshold callback function to the persistent data space map
interface for a subsequent patch to use.

dm-thin and dm-cache are interested in knowing when they're getting
low on metadata or data blocks. This patch introduces a new method
for registering a callback against a threshold.

Signed-off-by: Joe Thornber
Signed-off-by: Alasdair G Kergon

Joe Thornber
2013-05-10 21:37:20 +0800
24347e959 dm thin: detect metadata device resizing ... Browse Code »

Allow the dm thin pool metadata device to be extended.

Whenever a pool is resumed, detect whether the size of the metadata
device has increased, and if so, extend the metadata to use the new
space.

Signed-off-by: Joe Thornber
Signed-off-by: Alasdair G Kergon

Joe Thornber
2013-05-10 21:37:19 +0800
1921c56d9 dm persistent data: support space map resizing ... Browse Code »

Support extending a dm persistent data metadata space map.

The extend itself is implemented by switching back to the boostrap
allocator and pointing to the new space. The extra bitmap indexes are
then allocated from the new space, and finally we switch back to the
proper space map ops and tweak the reference counts.

Signed-off-by: Joe Thornber
Signed-off-by: Alasdair G Kergon

Joe Thornber
2013-05-10 21:37:19 +0800
5d0db96d1 dm thin: open dev read only when possible ... Browse Code »

If a thin pool is created in read-only-metadata mode then only open the
metadata device read-only.

Previously it was always opened with FMODE_READ | FMODE_WRITE.

(Note that dm_get_device() still allows read-only dm devices to be used
read-write at the moment: If I create a read-only linear device for the
metadata, via dmsetup load --readonly, then I can still create a rw pool
out of it.)

Signed-off-by: Joe Thornber
Signed-off-by: Alasdair G Kergon

Joe Thornber
2013-05-10 21:37:19 +0800
b17446df2 dm thin: refactor data dev resize ... Browse Code »

Refactor device size functions in preparation for similar metadata
device resizing functions.

Signed-off-by: Joe Thornber
Signed-off-by: Alasdair G Kergon

Joe Thornber
2013-05-10 21:37:18 +0800
8c5008fac dm cache: replace memcpy with struct assignment ... Browse Code »

Use struct assignment rather than memcpy in dm cache.

Signed-off-by: Joe Thornber
Signed-off-by: Alasdair G Kergon

Joe Thornber
2013-05-10 21:37:18 +0800
aeed1420a dm cache: fix typos in comments ... Browse Code »

Fix up some typos in dm-cache comments.

Signed-off-by: Joe Thornber
Signed-off-by: Alasdair G Kergon

Joe Thornber
2013-05-10 21:37:18 +0800
e12c1fd9d dm cache policy: fix description of lookup fn ... Browse Code »

Correct the documented requirement on the return code from dm cache policy
lookup functions stated in the policy module header file.

Signed-off-by: Alasdair G Kergon

Alasdair G Kergon
2013-05-10 21:37:17 +0800
88a488f62 dm persistent data: fix error message typos ... Browse Code »

Fix some typos in dm-space-map-metadata.c error messages.

Signed-off-by: Joe Thornber
Signed-off-by: Alasdair G Kergon

Joe Thornber
2013-05-10 21:37:17 +0800