HIGH LOAD AVERAGE BUT LOW CPU ULITIZATION IN LINUX

We have two node Oracle SOA cluster. Both nodes get almost same number of requests but we observed that on one node load average was way higher than second node though cpu utilization was almost same.

1) One reason could be for high load average is if system has lot of processes that are uninterruptible state. 

2) Checked if there were any processes in  ‘D’ (Disk wait) state.

top -b -n 1 | awk ‘{if (NR <=7) print; else if ($8 == “D”) print $1 ” ” $8 ” ”  $12 }’

Output:

PID USER      PR  NI  VIRT  RES  SHR S %CPU %MEM    TIME+  COMMAND
14688 D df
16092 D df
17324 D df
18610 D df
19192 D df
19735 D df
19759 D df
20619 D df

3) There are lot df processes waiting for disk. Find the time since they have been waiting for disk.

ps -eo pid,cmd,lstart | grep df
14688 df -h                       Tue May 16 22:45:01 2017
16092 df -h                       Tue May 16 22:50:01 2017
17324 df -h                       Tue May 16 22:55:01 2017
18610 df -h                       Tue May 16 23:00:01 2017
19192 df -kh                      Tue May 16 23:01:46 2017
19735 df -kh                      Tue May 16 23:04:54 2017
19759 df -h                       Tue May 16 23:05:01 2017

From above, we can conclude that df commands started on May 16th were hanging (possibly due to NFS issue.) Killing the processed reduced load average gradually.

Comments

Popular posts from this blog

HOW WE REDUCED SOA OSB PROVISIONING FROM 4 DAYS TO 4 HOURS

RABBITMQ CONNECTION ERROR: JAVAX.NET.SSL.SSLHANDSHAKEEXCEPTION: INVALID ECDH SERVERKEYEXCHANGE SIGNATURE

NOT ABLE TO START RABBITMQ CLUSTER: CANNOT DECLARE A QUEUE ‘~S’ ON NODE ‘~S’: ~255P