Posts

RABBITMQCTL LIST_QUEUES – BADRPC, SOME QUEUE(S) ARE UNRESPONSIVE

  We have a monitoring script that monitors rabbitmq queues and sends notification if message count is more than a threshold. Today the script stopped working. Followed below steps to troubleshoot issue: Logged in into first rabbitmq node and ran rabbitmqctl list_queues. It hung for 60 seconds and threw below output: rabbit@rmq-node1 ~]$ rabbitmqctl list_queues Timeout: 60.0 seconds … Listing queues for vhost / … name messages q1 0 q2 0 {:badrpc, {:timeout, 60.0, “Some queue(s) are unresponsive, use list_unresponsive_queues command.”}} As suggested by above output, ran  rabbitmqctl list_unresponsive_queues  . It too appeared hung. The cluster has 7 nodes. So, thought of checking if the command hangs in other nodes as well. Surprisingly it didn’t hang in one of the nodes. We hypothesized that this node has network connectivity issues with other nodes. The ping commands to other nodes were failing but pings from other nodes to this problematic node was succeeding. After inf...

SFTP ISSUES

  We have faced couple of sftp connectivity issues in our non prod environments in the past one week. Here is brief description about the issues and how we resolved them. Issue 1: sftp connection hangs without throwing any error. Public ip of one of the third party applications we are connecting got changed. We sent change in firewall rule to our firewall team. Team made change post that too we were not able to connect sftp server. Here is the verbose output: sftp -vvv -P 22222 user@104.x.x.x OpenSSH_7.4p1, OpenSSL 1.0.2k-fips 26 Jan 2017 debug1: Reading configuration data /etc/ssh/ssh_config debug1: /etc/ssh/ssh_config line 58: Applying options for * debug2: resolving “104.x.x.x” port 22222 debug2: ssh_connect_direct: needpriv 0 debug1: Connecting to 104.x.x.x [104.x.x.x] port 22222. debug1: Connection established. debug1: identity file /home/user/.ssh/id_rsa type 1 debug1: key_load_public: No such file or directory debug1: identity file /home/user/.ssh/id_rsa-cert type -1 de...

WEBLOGIC MANAGED SERVER START HANGING AFTER DEPLOYMENT

 We have a 4 node test SOA/Weblogic cluster. After CI deployment, the pipeline automatically restarted SOA managed servers. But managed servers in all 4 nodes stuck in STARTING state. We faced similar startup issues earlier. So, revisited these pages:  Issue 1 ,  Issue 2 ,  Issue 3 ,  Issue 4 , to check if the current issue is similar to earlier issues. Unfortunately, it is not related to any of them. Ran lsof ( Issue 3 ) output to check if server is waiting for any remote connection. But there isn’t any. Ran strace and thread dump but could not pick anything. As admin server is running without issue and problem is with only managed servers, I guessed that problem could be with databases that SOA adapters are connecting to. So, removed the targets from all datasources like EBS (excluded SOA DB as Admin server is up without any issue). Now, I was able to start managed servers. In order to identify the culprit, added targets to one datasource at a time and able to...

TARGET SERVERS MISSING FROM DATASOURCE TESTING TAB

When we get errors related to datasource in weblogic, first thing we do is go to datasource Monitoring -> Testing tab and test datasource. Sometimes, targets might be missing from this tab. During server startup, server tries to load datasource and create connections to database. If there is any underlying issue with database like password change, password expiry, unavailability of database, it doesn’t load the datasource and we see missing targets in testing tab. Some of the errors, we can see are: ####<Apr 21, 2022 3:44:12,251 AM PDT> <Info> <JDBC> <orasoa-test12-w2> <WLS_SOA2> <[ACTIVE] ExecuteThread: '0' for queue: 'weblogic.kernel.Default (self-tuning)'> <<WLS Kernel>> <> <d47e58ce-a023-4de2-acb7-24391c8056df-0000000a> <1650537852251> <[severity-value: 64] [rid: 0] [partition-id: 0] [partition-name: DOMAIN] > <BEA-001508> <Destroying data source connection pool TestDS.> ####<A...

RABBITMQ INTEGRATION WITH WAVEFRONT THROUGH TELEGRAF AGENT

  We wanted to monitor Rabbitmq through Wavefront. So, we followed below steps to configure the integration: Install telegraf agent. Enable Rabbitmq management plugin. Configure telegraf agent to use Rabbitmq input plugin. Restart telegraf Even after completing above steps we were not able to see complete data in Wavefront dashboards. We were able to see basic node level data but not able to see queue level data. We checked /var/log/messages where telegraf logs are written and found following messages. Apr 10 03:48:36 test-app1 telegraf: 2022-04-10T10:48:36Z E! [inputs.rabbitmq] Error in plugin: getting “/api/overview” failed: 401 Unauthorized Apr 10 03:48:36 test-app1 telegraf: 2022-04-10T10:48:36Z E! [inputs.rabbitmq] Error in plugin: getting “/api/nodes” failed: 401 Unauthorized Apr 10 03:48:36 test-app1 telegraf: 2022-04-10T10:48:36Z E! [inputs.rabbitmq] Error in plugin: getting “/api/exchanges” failed: 401 Unauthorized Apr 10 03:48:36 test-app1 telegraf: 2022-04-10T10:48:36Z E...

WAVEFRONT ALERT QUERIES FOR KUBERNETES MONITORING

Recently, we have setup monitoring of Kubernetes using wavefront. Here are some of the useful alerts and their queries: POD Memory Utilization: ts(“kubernetes.pod.memory.working_set”, namespace_name=”xyz”)/ts(“kubernetes.pod.memory.limit”, namespace_name=”xyz”) * 100 Kong POD CPU Utilization: ts(“kubernetes.pod.cpu.usage_rate”, namespace_name=”xyz”)/ts(“kubernetes.pod.cpu.limit”, namespace_name=”xyz”) * 100 Kong Replica Count Mismatch: ts(“kubernetes.deployment.desired_replicas”, namespace_name=”xyz”) – ts(“kubernetes.deployment.available_replicas” and namespace_name=”xyz”) New Pod Created/Pod Deleted: highpass(0, ts(“kubernetes.pod.uptime”, namespace_name=”xyz”) < 630000) Container Restart: mdiff(10m, ts(“kubernetes.pod.restart_count”, namespace_name=”xyz”))

SCRIPT WORKING PERFECTLY FROM COMMAND LINE BUT FAILING WITH CRONTAB

  I was writing small bash script to take backup of a application configuration. It was tested from command line and everything looks fine. Later, I had setup a cronjob to run it at a scheduled time. But there was no backup. One of the common reasons for this is when we use relative paths in the script. crontab default working directory is user’s home directory. If scripts have relative paths, it will not be able to find right files and directories. I checked if there were any such relative paths but nothing was found. To troubleshoot the issue, placed echo almost after every single statement. After analyzing these echo outputs, I suspected that crontab is not able to find one of the commands used in the script. But this utility was already installed and got confirmation by using “which” command. That means that PATH variable used by crontab is different from regular login shell. Got it confirmed by printing PATH variable. To fix the issue, modified PATH variable in the script to i...