Posts

KAFKA CONSUMER CLI ERROR: TIMED OUT WAITING FOR A NODE ASSIGNMENT

This post is more of a personal note than a perfect analysis. I have applied lot of workarounds and still not able to identify exactly what caused issue. Writing this blog so it will act as future reference if similar issue happens and add more details about the issue. We were adding monitoring for Kafka consumer lag using kafka-consumer-groups.sh which is located under Kafka installation bin directory. We were able to get lag in all environment except one and script execution failed with below error ( $KAFKA_BIN/kafka-consumer-groups.sh –describe –all-groups –bootstrap-server localhost:9092 ): java.util.concurrent.ExecutionException: org.apache.kafka.common.errors.TimeoutException: Timed out waiting for a node assignment. Call: metadata at java.util.concurrent.CompletableFuture.reportGet(CompletableFuture.java:357) at java.util.concurrent.CompletableFuture.get(CompletableFuture.java:1908) at org.apache.kafka.common.internals.KafkaFutureImpl.get(KafkaFutureImpl.java:165) at kafka.admin...

HOW TO FIND THE ENDPOINT A JAVA THREAD IS READING FROM

Image
  We may experience performance problems due to stuck threads i.e., when the thread is not making any progress and just stuck. Typically this may be due to contention with other threads, waiting for an event or waiting for data to be arrived on a network connection. This blog post is about finding other end of the connection which got stuck. If the application is simple and it is connecting to couple of hosts, it is very easy to identify other end of connection. If the application connects multiple boundary systems, databases, it becomes hard to identify the slow connection especially if the code path for all these end points is same. Middleware is one such application. Some applications have built in mechanism to identify this problem. For example, Oracle SOA and OSB generates diagnostic dumps under $DOMAIN_HOME/servers/<serverName>/adr/diag/ofm/<domainName>/<serverName>/incident. This directory contains readme.txt with details like composite, osb service that th...

RABBITMQCTL LIST_QUEUES – BADRPC, SOME QUEUE(S) ARE UNRESPONSIVE

  We have a monitoring script that monitors rabbitmq queues and sends notification if message count is more than a threshold. Today the script stopped working. Followed below steps to troubleshoot issue: Logged in into first rabbitmq node and ran rabbitmqctl list_queues. It hung for 60 seconds and threw below output: rabbit@rmq-node1 ~]$ rabbitmqctl list_queues Timeout: 60.0 seconds … Listing queues for vhost / … name messages q1 0 q2 0 {:badrpc, {:timeout, 60.0, “Some queue(s) are unresponsive, use list_unresponsive_queues command.”}} As suggested by above output, ran  rabbitmqctl list_unresponsive_queues  . It too appeared hung. The cluster has 7 nodes. So, thought of checking if the command hangs in other nodes as well. Surprisingly it didn’t hang in one of the nodes. We hypothesized that this node has network connectivity issues with other nodes. The ping commands to other nodes were failing but pings from other nodes to this problematic node was succeeding. After inf...

SFTP ISSUES

  We have faced couple of sftp connectivity issues in our non prod environments in the past one week. Here is brief description about the issues and how we resolved them. Issue 1: sftp connection hangs without throwing any error. Public ip of one of the third party applications we are connecting got changed. We sent change in firewall rule to our firewall team. Team made change post that too we were not able to connect sftp server. Here is the verbose output: sftp -vvv -P 22222 user@104.x.x.x OpenSSH_7.4p1, OpenSSL 1.0.2k-fips 26 Jan 2017 debug1: Reading configuration data /etc/ssh/ssh_config debug1: /etc/ssh/ssh_config line 58: Applying options for * debug2: resolving “104.x.x.x” port 22222 debug2: ssh_connect_direct: needpriv 0 debug1: Connecting to 104.x.x.x [104.x.x.x] port 22222. debug1: Connection established. debug1: identity file /home/user/.ssh/id_rsa type 1 debug1: key_load_public: No such file or directory debug1: identity file /home/user/.ssh/id_rsa-cert type -1 de...

WEBLOGIC MANAGED SERVER START HANGING AFTER DEPLOYMENT

 We have a 4 node test SOA/Weblogic cluster. After CI deployment, the pipeline automatically restarted SOA managed servers. But managed servers in all 4 nodes stuck in STARTING state. We faced similar startup issues earlier. So, revisited these pages:  Issue 1 ,  Issue 2 ,  Issue 3 ,  Issue 4 , to check if the current issue is similar to earlier issues. Unfortunately, it is not related to any of them. Ran lsof ( Issue 3 ) output to check if server is waiting for any remote connection. But there isn’t any. Ran strace and thread dump but could not pick anything. As admin server is running without issue and problem is with only managed servers, I guessed that problem could be with databases that SOA adapters are connecting to. So, removed the targets from all datasources like EBS (excluded SOA DB as Admin server is up without any issue). Now, I was able to start managed servers. In order to identify the culprit, added targets to one datasource at a time and able to...

TARGET SERVERS MISSING FROM DATASOURCE TESTING TAB

When we get errors related to datasource in weblogic, first thing we do is go to datasource Monitoring -> Testing tab and test datasource. Sometimes, targets might be missing from this tab. During server startup, server tries to load datasource and create connections to database. If there is any underlying issue with database like password change, password expiry, unavailability of database, it doesn’t load the datasource and we see missing targets in testing tab. Some of the errors, we can see are: ####<Apr 21, 2022 3:44:12,251 AM PDT> <Info> <JDBC> <orasoa-test12-w2> <WLS_SOA2> <[ACTIVE] ExecuteThread: '0' for queue: 'weblogic.kernel.Default (self-tuning)'> <<WLS Kernel>> <> <d47e58ce-a023-4de2-acb7-24391c8056df-0000000a> <1650537852251> <[severity-value: 64] [rid: 0] [partition-id: 0] [partition-name: DOMAIN] > <BEA-001508> <Destroying data source connection pool TestDS.> ####<A...

RABBITMQ INTEGRATION WITH WAVEFRONT THROUGH TELEGRAF AGENT

  We wanted to monitor Rabbitmq through Wavefront. So, we followed below steps to configure the integration: Install telegraf agent. Enable Rabbitmq management plugin. Configure telegraf agent to use Rabbitmq input plugin. Restart telegraf Even after completing above steps we were not able to see complete data in Wavefront dashboards. We were able to see basic node level data but not able to see queue level data. We checked /var/log/messages where telegraf logs are written and found following messages. Apr 10 03:48:36 test-app1 telegraf: 2022-04-10T10:48:36Z E! [inputs.rabbitmq] Error in plugin: getting “/api/overview” failed: 401 Unauthorized Apr 10 03:48:36 test-app1 telegraf: 2022-04-10T10:48:36Z E! [inputs.rabbitmq] Error in plugin: getting “/api/nodes” failed: 401 Unauthorized Apr 10 03:48:36 test-app1 telegraf: 2022-04-10T10:48:36Z E! [inputs.rabbitmq] Error in plugin: getting “/api/exchanges” failed: 401 Unauthorized Apr 10 03:48:36 test-app1 telegraf: 2022-04-10T10:48:36Z E...