Posts

WEBLOGIC MANAGED SERVER START HANGING AFTER DEPLOYMENT

 We have a 4 node test SOA/Weblogic cluster. After CI deployment, the pipeline automatically restarted SOA managed servers. But managed servers in all 4 nodes stuck in STARTING state. We faced similar startup issues earlier. So, revisited these pages:  Issue 1 ,  Issue 2 ,  Issue 3 ,  Issue 4 , to check if the current issue is similar to earlier issues. Unfortunately, it is not related to any of them. Ran lsof ( Issue 3 ) output to check if server is waiting for any remote connection. But there isn’t any. Ran strace and thread dump but could not pick anything. As admin server is running without issue and problem is with only managed servers, I guessed that problem could be with databases that SOA adapters are connecting to. So, removed the targets from all datasources like EBS (excluded SOA DB as Admin server is up without any issue). Now, I was able to start managed servers. In order to identify the culprit, added targets to one datasource at a time and able to...

TARGET SERVERS MISSING FROM DATASOURCE TESTING TAB

When we get errors related to datasource in weblogic, first thing we do is go to datasource Monitoring -> Testing tab and test datasource. Sometimes, targets might be missing from this tab. During server startup, server tries to load datasource and create connections to database. If there is any underlying issue with database like password change, password expiry, unavailability of database, it doesn’t load the datasource and we see missing targets in testing tab. Some of the errors, we can see are: ####<Apr 21, 2022 3:44:12,251 AM PDT> <Info> <JDBC> <orasoa-test12-w2> <WLS_SOA2> <[ACTIVE] ExecuteThread: '0' for queue: 'weblogic.kernel.Default (self-tuning)'> <<WLS Kernel>> <> <d47e58ce-a023-4de2-acb7-24391c8056df-0000000a> <1650537852251> <[severity-value: 64] [rid: 0] [partition-id: 0] [partition-name: DOMAIN] > <BEA-001508> <Destroying data source connection pool TestDS.> ####<A...

RABBITMQ INTEGRATION WITH WAVEFRONT THROUGH TELEGRAF AGENT

  We wanted to monitor Rabbitmq through Wavefront. So, we followed below steps to configure the integration: Install telegraf agent. Enable Rabbitmq management plugin. Configure telegraf agent to use Rabbitmq input plugin. Restart telegraf Even after completing above steps we were not able to see complete data in Wavefront dashboards. We were able to see basic node level data but not able to see queue level data. We checked /var/log/messages where telegraf logs are written and found following messages. Apr 10 03:48:36 test-app1 telegraf: 2022-04-10T10:48:36Z E! [inputs.rabbitmq] Error in plugin: getting “/api/overview” failed: 401 Unauthorized Apr 10 03:48:36 test-app1 telegraf: 2022-04-10T10:48:36Z E! [inputs.rabbitmq] Error in plugin: getting “/api/nodes” failed: 401 Unauthorized Apr 10 03:48:36 test-app1 telegraf: 2022-04-10T10:48:36Z E! [inputs.rabbitmq] Error in plugin: getting “/api/exchanges” failed: 401 Unauthorized Apr 10 03:48:36 test-app1 telegraf: 2022-04-10T10:48:36Z E...

WAVEFRONT ALERT QUERIES FOR KUBERNETES MONITORING

Recently, we have setup monitoring of Kubernetes using wavefront. Here are some of the useful alerts and their queries: POD Memory Utilization: ts(“kubernetes.pod.memory.working_set”, namespace_name=”xyz”)/ts(“kubernetes.pod.memory.limit”, namespace_name=”xyz”) * 100 Kong POD CPU Utilization: ts(“kubernetes.pod.cpu.usage_rate”, namespace_name=”xyz”)/ts(“kubernetes.pod.cpu.limit”, namespace_name=”xyz”) * 100 Kong Replica Count Mismatch: ts(“kubernetes.deployment.desired_replicas”, namespace_name=”xyz”) – ts(“kubernetes.deployment.available_replicas” and namespace_name=”xyz”) New Pod Created/Pod Deleted: highpass(0, ts(“kubernetes.pod.uptime”, namespace_name=”xyz”) < 630000) Container Restart: mdiff(10m, ts(“kubernetes.pod.restart_count”, namespace_name=”xyz”))

SCRIPT WORKING PERFECTLY FROM COMMAND LINE BUT FAILING WITH CRONTAB

  I was writing small bash script to take backup of a application configuration. It was tested from command line and everything looks fine. Later, I had setup a cronjob to run it at a scheduled time. But there was no backup. One of the common reasons for this is when we use relative paths in the script. crontab default working directory is user’s home directory. If scripts have relative paths, it will not be able to find right files and directories. I checked if there were any such relative paths but nothing was found. To troubleshoot the issue, placed echo almost after every single statement. After analyzing these echo outputs, I suspected that crontab is not able to find one of the commands used in the script. But this utility was already installed and got confirmation by using “which” command. That means that PATH variable used by crontab is different from regular login shell. Got it confirmed by printing PATH variable. To fix the issue, modified PATH variable in the script to i...

NOT ABLE TO START BOOMI ATOM AFTER JAVA UPDATE

We recently updated Boomi Java version to 11.0.14 through Boomi UI. When we were trying to restart atoms, one of the atoms refused to start. Unfortunately, no error message was written to logs. Luckily, there is a run option for atom script which usually throws error when there is an issue. I usually use this option for diagnosing startup issues. When executed ./atom run from command line, got below message: Error: Password file not found: $molecule_home/jre/lib/management/jmxremote.password Apparently, we forgot to backup and copy jmxremote.password file even though it was there in Boomi documentation. We copied jmxremote.password from earlier backups and able to start atom.

HOW TO FIND WHICH BOOMI PROCESS IS OCCUPYING DISK SPACE

  There could be multiple reasons for higher utilization of disk space in Boomi. Some of the reasons are too many files in following directories: logs – Stores atom logs execution – Logs of each process execution data – Data processed by atom processes processes – Deployed processes Some time ago, we received higher disk usage alert. We cleared some disk space by deleting files from logs and execution directories. Issue reoccurred after couple of days. We concluded that it was not regular disk fill-up and wanted to investigate further. We logged into terminal and issued following command from boomi molecule home directory to find out which directory was occupying more space. Found that execution directory was occupying more space. du -sh * | sort -rh Next, we wanted to check from which date this started occurring by using below command: cd $molecule_home/execution find . | awk -F'/' ' { print $2 } ' | uniq -c From output of the above command, we confirmed that issue wa...