Server bootet nicht mehr / Raid1

frannek

Registered User
Hallo,
ich habe gerade massivst ein Problem mit meinem Server. Erst dachte ich, dass mit grub was nicht stimmt aber jetzt sehe ich, dass im Raid was nicht okay ist und da komm ich nicht weiter:


mdadm --detail /dev/md0
/dev/md0:
Version : 1.0
Creation Time : Sun Sep 28 20:34:58 2014
Raid Level : raid1
Array Size : 524224 (512.02 MiB 536.81 MB)
Used Dev Size : 524224 (512.02 MiB 536.81 MB)
Raid Devices : 2
Total Devices : 1
Persistence : Superblock is persistent

Update Time : Mon Jul 18 11:42:44 2016
State : clean, degraded
Active Devices : 1
Working Devices : 1
Failed Devices : 0
Spare Devices : 0

Name : xxx.de:0
UUID : 5fa819f4:2c92ac34:cbc8bfd2:99d563a1
Events : 759

Number Major Minor RaidDevice State
0 0 0 0 removed
2 8 17 1 active sync /dev/sdb1




/dev/md1:
Version : 1.1
Creation Time : Sun Sep 28 20:35:00 2014
Raid Level : raid1
Array Size : 1463433024 (1395.64 GiB 1498.56 GB)
Used Dev Size : 1463433024 (1395.64 GiB 1498.56 GB)
Raid Devices : 2
Total Devices : 1
Persistence : Superblock is persistent

Intent Bitmap : Internal

Update Time : Mon Jul 18 11:38:27 2016
State : active, degraded
Active Devices : 1
Working Devices : 1
Failed Devices : 0
Spare Devices : 0

Name : xxx.de:1
UUID : 5b10b012:4b46eee3:65d4f8b6:0737b38e
Events : 38038187

Number Major Minor RaidDevice State
0 0 0 0 removed
2 8 19 1 active sync /dev/sdb3


cat /proc/mdstat
Personalities : [raid1] [raid0] [raid6] [raid5] [raid4]
md1 : active raid1 sdb3[2]
1463433024 blocks super 1.1 [2/1] [_U]
bitmap: 11/11 pages [44KB], 65536KB chunk

md0 : active raid1 sdb1[2]
524224 blocks super 1.0 [2/1] [_U]

unused devices: <none>


Kann mir hier jemand helfen? Ich sehe den Wald vor lauter Bäumen nicht mehr

Danke
 
Da fehlt eine Festplatte. Ob sie phys. fehlt oder "einfach tot" ist - keine Ahnung, das gibt für mich das Log nicht her.

-> Ticket beim Hoster aufmachen, er möge doch bitte (vermutlich) die erste HD tauschen.
 
Okay.. leuchtet ein, weil die zweite HDD vor paar Monaten ausgetauscht wurde. aber wieso bootet das System dann nicht mehr? Sollte es dennoch. Derzeit bin ich über ein Recovery System drin.
 
Last edited by a moderator:
ein Softwareraid bootet nur dann von allen beteiligten HDs, wenn auch der Bootloader auf alle HDs verteilt ist - da der in keinem mdm-Set enthalten ist muss da manuell nachgeholfen werden.
 
Also, die Platte scheint i.O.

smartctl -a /dev/sda
smartctl 5.41 2011-06-09 r3365 [x86_64-linux-3.13.0-61-generic] (local build)
Copyright (C) 2002-11 by Bruce Allen, http://smartmontools.sourceforge.net

=== START OF INFORMATION SECTION ===
Model Family: Seagate Barracuda LP
Device Model: ST31500541AS
Serial Number: 5XW2EVHH
LU WWN Device Id: 5 000c50 03254c3e2
Firmware Version: CC34
User Capacity: 1,500,301,910,016 bytes [1.50 TB]
Sector Size: 512 bytes logical/physical
Device is: In smartctl database [for details use: -P show]
ATA Version is: 8
ATA Standard is: ATA-8-ACS revision 4
Local Time is: Mon Jul 18 12:20:16 2016 UTC
SMART support is: Available - device has SMART capability.
SMART support is: Enabled

=== START OF READ SMART DATA SECTION ===
SMART overall-health self-assessment test result: PASSED

General SMART Values:
Offline data collection status: (0x82) Offline data collection activity
was completed without error.
Auto Offline Data Collection: Enabled.
Self-test execution status: ( 0) The previous self-test routine completed
without error or no self-test has ever
been run.
Total time to complete Offline
data collection: ( 653) seconds.
Offline data collection
capabilities: (0x7b) SMART execute Offline immediate.
Auto Offline data collection on/off support.
Suspend Offline collection upon new
command.
Offline surface scan supported.
Self-test supported.
Conveyance Self-test supported.
Selective Self-test supported.
SMART capabilities: (0x0003) Saves SMART data before entering
power-saving mode.
Supports SMART auto save timer.
Error logging capability: (0x01) Error logging supported.
General Purpose Logging supported.
Short self-test routine
recommended polling time: ( 1) minutes.
Extended self-test routine
recommended polling time: ( 255) minutes.
Conveyance self-test routine
recommended polling time: ( 2) minutes.
SCT capabilities: (0x103f) SCT Status supported.
SCT Error Recovery Control supported.
SCT Feature Control supported.
SCT Data Table supported.

SMART Attributes Data Structure revision number: 10
Vendor Specific SMART Attributes with Thresholds:
ID# ATTRIBUTE_NAME FLAG VALUE WORST THRESH TYPE UPDATED WHEN_FAILED RAW_VALUE
1 Raw_Read_Error_Rate 0x000f 112 099 006 Pre-fail Always - 46280712
3 Spin_Up_Time 0x0003 100 100 000 Pre-fail Always - 0
4 Start_Stop_Count 0x0032 100 100 020 Old_age Always - 21
5 Reallocated_Sector_Ct 0x0033 100 100 036 Pre-fail Always - 0
7 Seek_Error_Rate 0x000f 093 060 030 Pre-fail Always - 2462561024
9 Power_On_Hours 0x0032 046 046 000 Old_age Always - 47502
10 Spin_Retry_Count 0x0013 100 100 097 Pre-fail Always - 0
12 Power_Cycle_Count 0x0032 100 100 020 Old_age Always - 22
183 Runtime_Bad_Block 0x0032 095 095 000 Old_age Always - 5
184 End-to-End_Error 0x0032 100 100 099 Old_age Always - 0
187 Reported_Uncorrect 0x0032 100 100 000 Old_age Always - 0
188 Command_Timeout 0x0032 100 094 000 Old_age Always - 12885098643
189 High_Fly_Writes 0x003a 083 083 000 Old_age Always - 17
190 Airflow_Temperature_Cel 0x0022 059 052 045 Old_age Always - 41 (Min/Max 39/41)
194 Temperature_Celsius 0x0022 041 048 000 Old_age Always - 41 (0 22 0 0)
195 Hardware_ECC_Recovered 0x001a 052 029 000 Old_age Always - 46280712
197 Current_Pending_Sector 0x0012 098 098 000 Old_age Always - 87
198 Offline_Uncorrectable 0x0010 098 098 000 Old_age Offline - 87
199 UDMA_CRC_Error_Count 0x003e 200 200 000 Old_age Always - 0
240 Head_Flying_Hours 0x0000 100 253 000 Old_age Offline - 134557030463284
241 Total_LBAs_Written 0x0000 100 253 000 Old_age Offline - 312095054
242 Total_LBAs_Read 0x0000 100 253 000 Old_age Offline - 367176768

SMART Error Log Version: 1
No Errors Logged

SMART Self-test log structure revision number 1
No self-tests have been logged. [To run self-tests, use: smartctl -t]


SMART Selective self-test log data structure revision number 1
SPAN MIN_LBA MAX_LBA CURRENT_TEST_STATUS
1 0 0 Not_testing
2 0 0 Not_testing
3 0 0 Not_testing
4 0 0 Not_testing
5 0 0 Not_testing
Selective self-test flags (0x0):
After scanning selected spans, do NOT read-scan remainder of disk.
If Selective self-test is pending on power-up, resume after 0 minute delay.


Wie gehe ich nun vor? Ich hab keine Idee mehr
 
genau davor hab ich angst. Die kann ja nicht so von retzt auf gleich verschwinden. Also wirklich wie eine neue platte behandeln?
 
Bootet der Server wirklich nicht mehr oder geht es nur darum das RAID wieder zu rebuilden?

Wenn die alte HDD ok ist, dann sollten dort ja noch die Paritionen drauf sein:

raidhotadd /dev/md0 /dev/sda1
raidhotadd /dev/md1 /dev/sda3
 
noch mal geschaut:
Code:
 1 Raw_Read_Error_Rate 0x000f 112 099 006 Pre-fail Always - 46280712
 7 Seek_Error_Rate 0x000f 093 060 030 Pre-fail Always - 2462561024
sieht für mich nicht nach "die Platte ist ok" aus...
 
Stimmt, gut ist was anderes. Wundert mich, da smart sich normalerweise lauthals meldet. HDD wird ausgetauscht und dann hoffe ich, dass ich das mit Grub hinbekomme und damals bei der zweiten platte alles richtig machte.
 
So, neue platte ist drin und Grub2 bekmme ich noch immer nicht drauf. Die Partitionen sind angelegt und er ist auch kräftig am spiegeln. grub geht.. grub2 nicht und der server bootet leider auch nicht. Ich gehe nach dieser Manual vor:

http://adminforge.de/raid/mdadm/mdadm-raid-1-reparieren-nach-festplattentausch/
(jaja, link posten und so.. aber wie soll ichs sonst zeigen ;-)

Jemand noch ne Idee? Muss ich warten, bis die platte komplett gesynct wurde ?

Hier von Grub die Meldungen:

grub> root (hd0,0)
root (hd0,0)
Filesystem type is ext2fs, partition type 0x83
grub> setup (hd0)
setup (hd0)
Checking if "/boot/grub/stage1" exists... no
Checking if "/grub/stage1" exists... yes
Checking if "/grub/stage2" exists... yes
Checking if "/grub/e2fs_stage1_5" exists... yes
Running "embed /grub/e2fs_stage1_5 (hd0)"... failed (this is not fatal)
Running "embed /grub/e2fs_stage1_5 (hd0,0)"... failed (this is not fatal)
Running "install /grub/stage1 (hd0) /grub/stage2 p /grub/grub.conf "... succeeded
Done.
grub> setup (hd0)
setup (hd0)
Checking if "/boot/grub/stage1" exists... no
Checking if "/grub/stage1" exists... yes
Checking if "/grub/stage2" exists... yes
Checking if "/grub/e2fs_stage1_5" exists... yes
Running "embed /grub/e2fs_stage1_5 (hd0)"... failed (this is not fatal)
Running "embed /grub/e2fs_stage1_5 (hd0,0)"... failed (this is not fatal)
Running "install /grub/stage1 (hd0) /grub/stage2 p /grub/grub.conf "... succeeded
Done.
grub> setup (hd1)
setup (hd1)
Checking if "/boot/grub/stage1" exists... no
Checking if "/grub/stage1" exists... yes
Checking if "/grub/stage2" exists... yes
Checking if "/grub/e2fs_stage1_5" exists... yes
Running "embed /grub/e2fs_stage1_5 (hd1)"... failed (this is not fatal)
Running "embed /grub/e2fs_stage1_5 (hd0,0)"... failed (this is not fatal)
Running "install /grub/stage1 d (hd1) /grub/stage2 p /grub/grub.conf "... succeeded
Done.

Da kann ja was nicht stimmen


Grub2 meldet dies:

[root@zuluxxx /]# grub-install /dev/sdb
Could not find device for
[root@zuluxxx /]# grub-install /dev/sda
Could not find device for


grub.conf:


# grub.conf generated by anaconda
#
# Note that you do not have to rerun grub after making changes to this file
# NOTICE: You have a /boot partition. This means that
# all kernel and initrd paths are relative to /boot/, eg.
# root (hd0,0)
# kernel /vmlinuz-version ro root=/dev/md1
# initrd /initrd-[generic-]version.img
#boot=/dev/sda
default=0
timeout=5
splashimage=(hd0,0)/grub/splash.xpm.gz
hiddenmenu
title CentOS (2.6.32-642.1.1.el6.x86_64)
root (hd0,0)
kernel /vmlinuz-2.6.32-642.1.1.el6.x86_64 ro root=UUID=9be9622b-5faa-4db6-a255-9549bf51ea69 rd_NO_LUKS KEYBOARDTYPE=pc KEYTABLE=us LANG=de_DE.UTF-8 nodmraid SYSFONT=latarcyrheb-sun16 rd_MD_UUID=5b10b012:4b46eee3:65d4f8b6:0737b38e crashkernel=auto rd_NO_LVM rd_NO_DM rhgb quiet
initrd /initramfs-2.6.32-642.1.1.el6.x86_64.img
title CentOS (2.6.32-573.22.1.el6.x86_64)
root (hd0,0)
kernel /vmlinuz-2.6.32-573.22.1.el6.x86_64 ro root=UUID=9be9622b-5faa-4db6-a255-9549bf51ea69 rd_NO_LUKS KEYBOARDTYPE=pc KEYTABLE=us LANG=de_DE.UTF-8 nodmraid SYSFONT=latarcyrheb-sun16 rd_MD_UUID=5b10b012:4b46eee3:65d4f8b6:0737b38e crashkernel=auto rd_NO_LVM rd_NO_DM rhgb quiet
initrd /initramfs-2.6.32-573.22.1.el6.x86_64.img
title CentOS (2.6.32-573.18.1.el6.x86_64)
root (hd0,0)
kernel /vmlinuz-2.6.32-573.18.1.el6.x86_64 ro root=UUID=9be9622b-5faa-4db6-a255-9549bf51ea69 rd_NO_LUKS KEYBOARDTYPE=pc KEYTABLE=us LANG=de_DE.UTF-8 nodmraid SYSFONT=latarcyrheb-sun16 rd_MD_UUID=5b10b012:4b46eee3:65d4f8b6:0737b38e crashkernel=auto rd_NO_LVM rd_NO_DM rhgb quiet
initrd /initramfs-2.6.32-573.18.1.el6.x86_64.img
title CentOS (2.6.32-573.12.1.el6.x86_64)
root (hd0,0)
kernel /vmlinuz-2.6.32-573.12.1.el6.x86_64 ro root=UUID=9be9622b-5faa-4db6-a255-9549bf51ea69 rd_NO_LUKS KEYBOARDTYPE=pc KEYTABLE=us LANG=de_DE.UTF-8 nodmraid SYSFONT=latarcyrheb-sun16 rd_MD_UUID=5b10b012:4b46eee3:65d4f8b6:0737b38e crashkernel=auto rd_NO_LVM rd_NO_DM rhgb quiet
initrd /initramfs-2.6.32-573.12.1.el6.x86_64.img
title CentOS (2.6.32-504.1.3.el6.x86_64)
root (hd0,0)
kernel /vmlinuz-2.6.32-504.1.3.el6.x86_64 ro root=UUID=9be9622b-5faa-4db6-a255-9549bf51ea69 rd_NO_LUKS KEYBOARDTYPE=pc KEYTABLE=us LANG=de_DE.UTF-8 nodmraid SYSFONT=latarcyrheb-sun16 rd_MD_UUID=5b10b012:4b46eee3:65d4f8b6:0737b38e crashkernel=auto rd_NO_LVM rd_NO_DM rhgb quiet
initrd /initramfs-2.6.32-504.1.3.el6.x86_64.img

Die Files liegen alle unter /boot/. Sind also da.. aber dennoch klappts nicht. Ich habs mal mit Webmin probiert:

GNU GRUB version 0.97 (640K lower / 3072K upper memory)

[ Minimal BASH-like line editing is supported. For the first word, TAB
lists possible command completions. Anywhere else TAB lists the possible
completions of a device/filename.]
grub> find /boot/grub/grub.conf

Error 15: File not found
grub>
 
Last edited by a moderator:
Back
Top