Posts

INTRIGUING 502 BAD GATEWAY ERROR WHILE INVOKING KONG THAT PROXIES BOOMI SERVICE

  Our dev team has exposed a Boomi service through Kong Gateway to be consumed by other teams. The service was working for most of the time but throwing 502 errors randomly. Initially, we thought it was network glitch and ignored those errors. As the testing progressed to higher environments, the error rate increased. Though it was below 1 %, due to limitations at both client and upstream we were tasked to analyze and fix the issue. The high level request flow is like this:   Client -> Kong Gateway -> Boomi VIP -> Boomi Atoms -> Upstream/backend service. We started analysis with Kong and Boomi logs. Found below entries: Kong : {“date”:1.664432446599347E9,”log”:”2022-09-29T06:20:46.599260909Z stderr F 2022/09/29 06:20:46 [error] 38#0: *1693441  upstream prematurely closed connection  while reading response header from upstream, client: 10.200.63.6, server: kong, request: \”POST /test/outbound/consumer HTTP/1.0\”, upstream: Boomi : 2022_09_28.shared_http_se...

TROUBLESHOOTING HASHICORP VAULT KUBERNETES AUTH ERROR

  We were trying to integrate Spring Cloud Gateway (SCG) on Kubernetes with HashiCorp. We followed the steps mentioned in vault   documentation . We were able to bring up vault and vault injector but SCG pods were stuck in init state. Found following error in SCG application pod vault-agent-init container logs: NAME READY STATUS RESTARTS AGE scg-0 0/2 Init:0/1 0 9h vault-0 1/1 Running 0 10h vault-agent-injector-5c89c7dfc5-n2v6v 1/1 Running 0 20h Cgo: disabled Log Level: info Version: Vault v1.8.4 Version Sha: 925bc650ad1d997e84fbb832f302a6bfe0105bbb 2022-09-30T16:24:28.007Z [INFO] sink.server: starting sink server 2022-09-30T16:24:28.007Z INFO creating watcher 2022-09-30T16:25:28.008Z [ERROR] auth.handler: error authenticating: error="context deadline exceeded" backoff=1s To verify that vault injector was connecting to right vault instance, we verified init agent config using below kubectl command: kubectl exec -it scg-0 -c va...

CONNECTION RESET ERROR – BOOMI CONNECTING TO SFTP SERVER

Image
  This blog is about connection reset error we received while connecting from Boomi atom to remote sftp. This issue can occur for any client connecting to a remote sftp server. The sftp server we were trying to connect was old sftp server and it has been used in many integrations. We haven’t faced reset issue for other integrations flow except this new one. So, we first suspected that there was something wrong with the process itself. We looked closely at the process. The process has multiple processes – one main process and multiple child processes. Main process checks for availability of files and child processes transfer files to Boomi atoms for further processing. The first main process and few child processes succeed always but last few process would fail always (usually after processing 7 or 8 files). We tweaked the process to drill down the issue. Here are few thing we tried: Disabled all other processes that connect to the same sftp server. Issue still occurred. We checked ...

HOW TO MONITOR KAFKA CONSUMER LAG

There are couple of ways to monitor Kafka consumer Lag. There is an open source project called   Apache Burrow   which has lot of features and easy to set up. It filters out lot of false positives too. If there are an limitations of using it in your project, we can write bash scripts using Kafka consumer utility which ships with Kafka. We have written one such script. If you are interested,   please check it out here .

KAFKA CONSUMER CLI ERROR: TIMED OUT WAITING FOR A NODE ASSIGNMENT

This post is more of a personal note than a perfect analysis. I have applied lot of workarounds and still not able to identify exactly what caused issue. Writing this blog so it will act as future reference if similar issue happens and add more details about the issue. We were adding monitoring for Kafka consumer lag using kafka-consumer-groups.sh which is located under Kafka installation bin directory. We were able to get lag in all environment except one and script execution failed with below error ( $KAFKA_BIN/kafka-consumer-groups.sh –describe –all-groups –bootstrap-server localhost:9092 ): java.util.concurrent.ExecutionException: org.apache.kafka.common.errors.TimeoutException: Timed out waiting for a node assignment. Call: metadata at java.util.concurrent.CompletableFuture.reportGet(CompletableFuture.java:357) at java.util.concurrent.CompletableFuture.get(CompletableFuture.java:1908) at org.apache.kafka.common.internals.KafkaFutureImpl.get(KafkaFutureImpl.java:165) at kafka.admin...

HOW TO FIND THE ENDPOINT A JAVA THREAD IS READING FROM

Image
  We may experience performance problems due to stuck threads i.e., when the thread is not making any progress and just stuck. Typically this may be due to contention with other threads, waiting for an event or waiting for data to be arrived on a network connection. This blog post is about finding other end of the connection which got stuck. If the application is simple and it is connecting to couple of hosts, it is very easy to identify other end of connection. If the application connects multiple boundary systems, databases, it becomes hard to identify the slow connection especially if the code path for all these end points is same. Middleware is one such application. Some applications have built in mechanism to identify this problem. For example, Oracle SOA and OSB generates diagnostic dumps under $DOMAIN_HOME/servers/<serverName>/adr/diag/ofm/<domainName>/<serverName>/incident. This directory contains readme.txt with details like composite, osb service that th...

RABBITMQCTL LIST_QUEUES – BADRPC, SOME QUEUE(S) ARE UNRESPONSIVE

  We have a monitoring script that monitors rabbitmq queues and sends notification if message count is more than a threshold. Today the script stopped working. Followed below steps to troubleshoot issue: Logged in into first rabbitmq node and ran rabbitmqctl list_queues. It hung for 60 seconds and threw below output: rabbit@rmq-node1 ~]$ rabbitmqctl list_queues Timeout: 60.0 seconds … Listing queues for vhost / … name messages q1 0 q2 0 {:badrpc, {:timeout, 60.0, “Some queue(s) are unresponsive, use list_unresponsive_queues command.”}} As suggested by above output, ran  rabbitmqctl list_unresponsive_queues  . It too appeared hung. The cluster has 7 nodes. So, thought of checking if the command hangs in other nodes as well. Surprisingly it didn’t hang in one of the nodes. We hypothesized that this node has network connectivity issues with other nodes. The ping commands to other nodes were failing but pings from other nodes to this problematic node was succeeding. After inf...