INTRIGUING 502 BAD GATEWAY ERROR WHILE INVOKING KONG THAT PROXIES BOOMI SERVICE

 Our dev team has exposed a Boomi service through Kong Gateway to be consumed by other teams. The service was working for most of the time but throwing 502 errors randomly. Initially, we thought it was network glitch and ignored those errors. As the testing progressed to higher environments, the error rate increased. Though it was below 1 %, due to limitations at both client and upstream we were tasked to analyze and fix the issue. The high level request flow is like this: Client -> Kong Gateway -> Boomi VIP -> Boomi Atoms -> Upstream/backend service.

We started analysis with Kong and Boomi logs. Found below entries:

Kong:

{“date”:1.664432446599347E9,”log”:”2022-09-29T06:20:46.599260909Z stderr F 2022/09/29 06:20:46 [error] 38#0: *1693441 upstream prematurely closed connection while reading response header from upstream, client: 10.200.63.6, server: kong, request: \”POST /test/outbound/consumer HTTP/1.0\”, upstream:

Boomi:

2022_09_28.shared_http_server.: [28/Sep/2022:23:00:34 -0700] “POST /ws/rest/party/consumer HTTP/1.1″ 499 0 “-” “Java/11.0.5” “7d09643c-f584-4435-8d40-31cbbb820f82” “execution-f18712ff-5524-4e47-8c30-d19d43e6c88f-2022.09.28” 58043

So, Kong which is client got premature closing of connection error. At the same time, Boomi which is server got “client closed connection” (http 499) error. So, we thought someone sitting in between client and server closed connection as both Kong and Boomi thought other side of the connection closed it. As network loadbalaner is connecting both client and server, we checked with our network team, if there was any issue with loadbalancer. Dev teams, network teams and other infrastructure teams huddled into a call. Network team verified LB and confirmed that there was no issue with LB. We also changed Kong upstream endpoint to one of the Boomi backend servers instead of LB. Still we got 502 error. So, we confirmed that it was not due to LB.

As network is out of question, Dev team sent some artificial load to the service, while we were monitoring systems. Someone from Dev team noticed that service was failing when Boomi was taking more than 30 seconds to send response back to Kong. With this data, we searched internet and found this link. As we are not sure if this could resolve our problem, someone suggested to decrease timeout of value to 10 seconds to replicate issue easily. So we changed com.boomi.container.sharedServer.http.maxIdleTime value in container.properties file to 10 seconds. That suggestion worked and we could see much higher failures when timeout was set to 10 seconds. At this point, we were confident about the change and increase timeout to much higher value. We monitored for failures for next couple of days and issue didn’t occur again. If Boomi service handler had honored timeout setout at webserver Jetty layer, this issue would have been identified much quickly.

Without the team effort, the issue would have taken much more time to resolve. Thanks for Aditya, Minu, Parvathy, Kalyan and Abhilash in helping me in learning new things.

Comments

Popular posts from this blog

HOW WE REDUCED SOA OSB PROVISIONING FROM 4 DAYS TO 4 HOURS

NOT ABLE TO START RABBITMQ CLUSTER: CANNOT DECLARE A QUEUE ‘~S’ ON NODE ‘~S’: ~255P

SOA SUITE 12.2.1.4 INSTALLATION: GOT EXCEPTION WHEN AUTO CONFIGURING THE SCHEMA COMPONENT(S) WITH DATA OBTAINED FROM SHADOW TABLE