Posts

WAVEFRONT ALERT QUERIES FOR KUBERNETES MONITORING

Recently, we have setup monitoring of Kubernetes using wavefront. Here are some of the useful alerts and their queries: POD Memory Utilization: ts(“kubernetes.pod.memory.working_set”, namespace_name=”xyz”)/ts(“kubernetes.pod.memory.limit”, namespace_name=”xyz”) * 100 Kong POD CPU Utilization: ts(“kubernetes.pod.cpu.usage_rate”, namespace_name=”xyz”)/ts(“kubernetes.pod.cpu.limit”, namespace_name=”xyz”) * 100 Kong Replica Count Mismatch: ts(“kubernetes.deployment.desired_replicas”, namespace_name=”xyz”) – ts(“kubernetes.deployment.available_replicas” and namespace_name=”xyz”) New Pod Created/Pod Deleted: highpass(0, ts(“kubernetes.pod.uptime”, namespace_name=”xyz”) < 630000) Container Restart: mdiff(10m, ts(“kubernetes.pod.restart_count”, namespace_name=”xyz”))

SCRIPT WORKING PERFECTLY FROM COMMAND LINE BUT FAILING WITH CRONTAB

  I was writing small bash script to take backup of a application configuration. It was tested from command line and everything looks fine. Later, I had setup a cronjob to run it at a scheduled time. But there was no backup. One of the common reasons for this is when we use relative paths in the script. crontab default working directory is user’s home directory. If scripts have relative paths, it will not be able to find right files and directories. I checked if there were any such relative paths but nothing was found. To troubleshoot the issue, placed echo almost after every single statement. After analyzing these echo outputs, I suspected that crontab is not able to find one of the commands used in the script. But this utility was already installed and got confirmation by using “which” command. That means that PATH variable used by crontab is different from regular login shell. Got it confirmed by printing PATH variable. To fix the issue, modified PATH variable in the script to i...

NOT ABLE TO START BOOMI ATOM AFTER JAVA UPDATE

We recently updated Boomi Java version to 11.0.14 through Boomi UI. When we were trying to restart atoms, one of the atoms refused to start. Unfortunately, no error message was written to logs. Luckily, there is a run option for atom script which usually throws error when there is an issue. I usually use this option for diagnosing startup issues. When executed ./atom run from command line, got below message: Error: Password file not found: $molecule_home/jre/lib/management/jmxremote.password Apparently, we forgot to backup and copy jmxremote.password file even though it was there in Boomi documentation. We copied jmxremote.password from earlier backups and able to start atom.

HOW TO FIND WHICH BOOMI PROCESS IS OCCUPYING DISK SPACE

  There could be multiple reasons for higher utilization of disk space in Boomi. Some of the reasons are too many files in following directories: logs – Stores atom logs execution – Logs of each process execution data – Data processed by atom processes processes – Deployed processes Some time ago, we received higher disk usage alert. We cleared some disk space by deleting files from logs and execution directories. Issue reoccurred after couple of days. We concluded that it was not regular disk fill-up and wanted to investigate further. We logged into terminal and issued following command from boomi molecule home directory to find out which directory was occupying more space. Found that execution directory was occupying more space. du -sh * | sort -rh Next, we wanted to check from which date this started occurring by using below command: cd $molecule_home/execution find . | awk -F'/' ' { print $2 } ' | uniq -c From output of the above command, we confirmed that issue wa...

UNABLE TO STOP BOOMI ATOM

We were updating Java version using Atmosphere UI and tried to restart the atoms from terminal. All atoms were successfully restarted except one atom. We were getting timeout error while running ./atom stop command. We could see following error message on cluster status page of atmosphere UI. EW_ID_MISMATCH There are differing view IDs in the various view snapshot files. Severe These nodes are reporting the issue: Node 1 10_x_x_x Found different view id [10_x_x_x|90] in "node.10_x_x_x.dat" Node 2 10_x_x_x Found different view id [10_x_x_x|101] in "node.10_x_x_x.dat" We analyzed atom logs but could not find the reason for timeout. We killed the process and removed above dat files from $MOLECULE_HOME/bin/views directory. But we got again same problem, we couldn’t stop the atom. So, to troubleshoot further, captured system calls using strace by attaching to running process. Following lines were repeating in the trace log. open(" /tmp/i4jdaemon__data_boomi_Boomi_At...

WEBLOGIC ADMIN SERVER STARTUP FAILING WITH FORCE_SHUTTING_DOWN

  We shutdown our test weblogic admin server to apply recent security patches. Our colleague applied the path and tried to start Admin server. Server didn’t come-up and was failing with below error: A MultiException has 6 exceptions. They are: weblogic.security.SecurityInitializationException: Authentication denied: Boot identity not valid. The user name or password or both from the boot identity file (boot.properties) is not valid. The boot identity may have been changed since the boot identity file was created. Please edit and update the boot identity file with the proper values of username and password. The first time the updated boot identity file is used to start the server, these new values are encrypted. java.lang.IllegalStateException: Unable to perform operation: post construct on weblogic.security.SecurityService java.lang.IllegalArgumentException: While attempting to resolve the dependencies of weblogic.jndi.internal.RemoteNamingService errors were found. Server state ch...

GREP COMMAND IS NOT RESPONDING AGAINST EXTREMELY SMALL AMOUNT OF DATA

  As a sysadmin we run grep command innumerable number of times every day. Today I got to search for some string in a directory. As usual I fired grep command. The command is similar to   grep -R hello *   . I didn’t get result immediately. Assumed that size of the directory was huge. Instead of checking further, I grabbed cup of tea and was waiting for results. While I emptying the tea cup, I was hoping that empty screen will be filled up with results. Was I too optimistic or finding reason for getting additional cup of tea? But it didn’t happen. So, I cancelled the command and checked the size of the directory. Surprisingly size was few MBs. So, it was time to pull out one my favorite debugging tools: strace. Fired the command as below: strace grep -R hello * Condensed output: execve(“/usr/bin/grep”, [“grep”, “-R”, “hello”, “-“, “1.sh”, “1.txt”, “1.yaml”, “2.txt”, “2.yaml”, “3.txt”, “cookies.txt”,, “deck_1.9.0_linux_amd64.tar.gz”, “dummy-key.pem”, …], 0x7ffe9be4ac08 /* ...