Server gefreezed

Der schreibt alle 2 Sekunden in eine Datei die in /root/ liegt.
Ich würde es so machen:
Code:
while [ 1 ]
do
ps aux > /root/procloc && date >> /root/proclog
done
Dann hast Du nur den letzten Record, wann der Server gecrasht ist :D
 
kann es sein, daß du so gar keine Ahnung von Linux/Unix hast? Und evtl. auch kein Bock, ein wenig selbst zu suchen und zu probieren?

probier mal ein
Code:
screen
 
screen ist mir schon klar habe ich auch bereits versucht. Aber ich bin es gewölhnt sceens einen namen zu geben z.B. log worüber ich den screen später wieder aufrufen kann. Leider funzt da aber mein Befehl: screen -dmS Log <befehl> nicht. Baer danke für deine Meinung bezüglich keinen Bocks...

Gruß Tim
 
Hallo.

Da mein Server soeben wieder gecrashed ist, aber in der proclog nichts verdächtiges drinne steh, habe ich auch hier nochmal die syslog und ähnliches gesendet. Ich hoffe ihr könnt mir helfen.

Gruß Tim

syslog:
Code:
Mar  6 10:00:30 s15430026 postfix/local[7497]: BC5A6E0AC6D: to=<[email protected]>, orig_to=<[email protected]>, relay=local, delay=0.05, delays=0.02/0/0/0.03, dsn=2.0.0, status=sent (forwarded as C1063E0AC6C)
Mar  6 10:00:30 s15430026 postfix/qmgr[2518]: C1063E0AC6C: from=<>, size=3451, nrcpt=1 (queue active)
Mar  6 10:00:30 s15430026 postfix/qmgr[2518]: BC5A6E0AC6D: removed
Mar  6 10:00:30 s15430026 postfix/smtp[7498]: C1063E0AC6C: to=<[email protected]>, orig_to=<[email protected]>, relay=none, delay=0.03, delays=0.03/0/0/0, dsn=5.4.4, status=bounced (Host or domain name not found. Name service error for name=timliebetra21u.de type=AAAA: Host not found)
Mar  6 10:00:30 s15430026 postfix/qmgr[2518]: C1063E0AC6C: removed
Mar  6 10:09:01 s15430026 /USR/SBIN/CRON[7511]: (root) CMD (  [ -x /usr/lib/php5/maxlifetime ] && [ -d /var/lib/php5 ] && find /var/lib/php5/ -type f -cmin +$(/usr/lib/php5/maxlifetime) -delete)
Mar  6 10:10:01 s15430026 /USR/SBIN/CRON[7521]: (root) CMD ([ -x /opt/psa/admin/sbin/backupmng ] && /opt/psa/admin/sbin/backupmng >/dev/null 2>&1)
Mar  6 10:17:01 s15430026 /USR/SBIN/CRON[7530]: (root) CMD (   cd / && run-parts --report /etc/cron.hourly)
Mar  6 10:25:01 s15430026 /USR/SBIN/CRON[7542]: (root) CMD ([ -x /opt/psa/admin/sbin/backupmng ] && /opt/psa/admin/sbin/backupmng >/dev/null 2>&1)
Mar  6 10:30:01 s15430026 /USR/SBIN/CRON[7553]: (drweb) CMD (/opt/drweb/update.pl)
Mar  6 10:30:01 s15430026 /USR/SBIN/CRON[7554]: (www-data) CMD (php5 -f /var/shoutcast/server.php)
Mar  6 10:30:01 s15430026 update.pl[7555]: Dr.Web (R) Updater ($Revision: 1.7.2.32.2.5 $) started ...
Mar  6 10:30:01 s15430026 update.pl[7555]: You appear to use some linux distro, but there is no lsb_release command accessible in your system. Either your PATH environment variable is screwed up (due to security considerations, or whatever), or your linux distribution does not fully support Linux Standard Base (proposed by Free Standards Group). It is probably okay. Falling back to another way of discovering the name of your linux distro.
Mar  6 10:30:01 s15430026 update.pl[7555]: You appear to use Debian GNU/Linux (the file '/etc/debian_version' exists, which is typical for Debian GNU/Linux). Trying to discover the exact version and hardware architecture of your linux system.
Mar  6 10:30:01 s15430026 update.pl[7555]: Path to bases      : /var/drweb/bases/
Mar  6 10:30:01 s15430026 update.pl[7555]: Path to URL list   : /var/drweb/bases/
Mar  6 10:30:01 s15430026 update.pl[7555]: Path to blacklists : /var/drweb/dws/
Mar  6 10:30:01 s15430026 update.pl[7555]: Path to lzma: /opt/drweb/lzma
Mar  6 10:30:01 s15430026 update.pl[7555]: custom URL list isn't defined
Mar  6 10:30:01 s15430026 update.pl[7555]: try using Dr.Web URL list (/var/drweb/bases/update.drl)
Mar  6 10:30:01 s15430026 update.pl[7555]: exec(/opt/drweb/read_signed drl /var/drweb/bases/update.drl) ...
Mar  6 10:30:01 s15430026 update.pl[7555]: no custom update servers
Mar  6 10:30:01 s15430026 update.pl[7555]: main update servers: http://update.drweb.com/unix/500, http://update.msk.drweb.com/unix/500, http://update.msk3.drweb.com/unix/500, http://update.us.drweb.com/unix/500, http://update.msk5.drweb.com/unix/500, http://update.msk6.drweb.com/unix/500, http://update.fr1.drweb.com/unix/500, http://update.us1.drweb.com/unix/500, http://update.nsk1.drweb.com/unix/500
Mar  6 10:30:01 s15430026 update.pl[7555]: drldir not found: "/var/drweb/drl", assuming there are no plugins to update
Mar  6 10:30:01 s15430026 postfix/pickup[7396]: C77D8E0AC6D: uid=33 from=<www-data>
Mar  6 10:30:01 s15430026 postfix/cleanup[7564]: C77D8E0AC6D: message-id=<[email protected]>
Mar  6 10:30:01 s15430026 update.pl[7555]: Attempting to fetch http://update.drweb.com/unix/500/drweb32.lst ...
Mar  6 10:30:01 s15430026 postfix/qmgr[2518]: C77D8E0AC6D: from=<[email protected]>, size=755, nrcpt=1 (queue active)
Mar  6 10:30:01 s15430026 postfix/error[7568]: C77D8E0AC6D: to=<[email protected]>, orig_to=<www-data>, relay=none, delay=0.11, delays=0.08/0/0/0.02, dsn=5.0.0, status=bounced (User unknown in virtual alias table)
Mar  6 10:30:01 s15430026 postfix/cleanup[7564]: D6833E0AC6C: message-id=<[email protected]>
Mar  6 10:30:01 s15430026 postfix/bounce[7569]: C77D8E0AC6D: sender non-delivery notification: D6833E0AC6C
Mar  6 10:30:01 s15430026 postfix/qmgr[2518]: D6833E0AC6C: from=<>, size=2853, nrcpt=1 (queue active)
Mar  6 10:30:01 s15430026 postfix/qmgr[2518]: C77D8E0AC6D: removed
Mar  6 10:30:01 s15430026 postfix/error[7568]: D6833E0AC6C: to=<[email protected]>, relay=none, delay=0.05, delays=0.02/0/0/0.03, dsn=5.0.0, status=bounced (User unknown in virtual alias table)
Mar  6 10:30:01 s15430026 postfix/qmgr[2518]: D6833E0AC6C: removed
Mar  6 10:30:03 s15430026 update.pl[7555]: request with 309 bytes length was sent to update.drweb.com
Mar  6 10:30:03 s15430026 update.pl[7555]: update.drweb.com return 200 OK
Mar  6 10:30:03 s15430026 update.pl[7555]: 4442 bytes received from http://update.drweb.com/unix/500/drweb32.lst.
Mar  6 10:30:03 s15430026 update.pl[7555]: downloading notifications ...
Mar  6 10:30:03 s15430026 update.pl[7555]: downloading updated files ...
Mar  6 10:30:03 s15430026 update.pl[7555]: downloading new/updated files ...
 

Attachments

tar cfvz /var/log.tar.gz /var/log

Die Datei dann hier hochladen und eine genaue Uhrzeit mit Datum nennen, wann der Server gecrasht ist. In der dmesg-Logfile steht nämlich 10:30 Uhr als letzter Log-Eintrag und es ist inzwischen 16:44. Damit kann hier keiner was anfangen. Um Dir helfen zu können, benötigen wir auch Deine Hilfe ;-)
 
Oki ist gemacht :). Die Uhrzeit ist zwischen 10:00 Uhr und 11 Uhr das ganze am 7. März. Eine genaue Zeit weis ich leider nicht - sry.

Ich habe inzwischen auch die CPU und Festplatten Temperatur gemessen die ist völlig ok. Desweiteren habe ich auch bei meinen Hoster nen Memtest und Festplatten sowie Netzteiltest in Auftrag gegeben.

Gruß Tim

DATEI-LINK: http://www.mediafire.com/?q0t6y9uvb7un2ag
 
Last edited by a moderator:
/var/log/syslog.0 fehlt/ist verschwunden. Ohne die, kann ich leider nichts anfangen, da Logeinträge zwischen
Code:
Mar  6 17:40:01
und
Code:
Mar  7 11:36:09
fehlen.

/Edit: Scheint wohl eher ein schwerwiegenderes Problem zu sein, da in jeder Log-File zwischen den beiden o.g. Zeiten kein Eintrag existiert.
Ich tippe da 'mal blind darauf, dass entweder das Hardware-Raid oder die Festplatten problematische Macken haben.
Sag' deinem Hoster 'mal bitte, dass er nichts, außer die Festplatte checken soll. Die schreibt nämlich nach einer unbestimmten Zeit nichts mehr und so crasht dann der Server.
 
Last edited by a moderator:
Bitte mal folgende Ausgaben posten:
Code:
du -h -d 1 /
df -h

Schon die S.M.A.R.T.-Werte kontrollliert und fsck laufen lassen?
 
Den Befehl kennt er nicht. Aber beim letzten Absturtz kam jetz folgendes:

BUG: unable to handle kernel NULL pointer dereference at (null)
[ *722.500915] IP: [<ffffffff811d8c2f>] acpi_ex_resolve_to_value+0x1f7/0x20c
[ *722.500915] PGD 100c3a067 PUD 11cd9a067 PMD 0*
[ *722.500915] Oops: 0002 [#1] SMP*

Kann jemand was damit anfangen? Oder hat wer ne Lösung?

Gruß Tim
 
Wichtig: Bitte einen Bugreport an die Kernel-Entwickler verfassen.

Workaround: ACPI per Bootparameter deaktivieren.
 
so danke dir. Habe es gemeldet und auch abgeschaltet. ich hoffe daran lag es warum der server immer abgestürzt ist.

Gruß Tim
 
So nun ist er wieder abgestürzt. Folgende Meldung kam:

ata2: SATA link up 3.0 Gbps (SStatus 123 SControl 300)
[ * *4.180431] ata2.00: ATA-8: ST500DM002-1BC142, JC4B, max UDMA/133
[ * *4.192610] ata2.00: 976773168 sectors, multi 16: LBA48 NCQ (depth 31/32)

Weis einer was das wieder ist?
Kann das mit dem Kernel zutun haben?
 
Last edited by a moderator:
Das sieht nach einem ganz normalen Status Eintrag aus, ich denke mal das hat keinen Einfluss auf die Abstürze.
 
Back
Top