MYSTERIOUS OHS HTTP 400 RESPONSE FINALLY SOLVED – OHS DELETING TMPDIR

 This story goes back to Aug’20 when we got a new engagement with a client to manage their infrastructure. This infrastructure had been managed by different vendor on their platform. So, we had to make few changes like DNS server, network domain. We have implemented domain name changes on a Saturday for all SOA/weblogic application and everything went fine.

Next Monday, business started using application hosted on weblogic fronteneded by OHS (collocated), users started getting 400 response code frequently. It was not 100% failure but some requests were going through. Someone in our team suggested restart of OHS. I tried to login to em console to restart OHS but log in was also failing with same 400 response code. I tried to restart using startComponent.sh script. It was failed due to misconfiguration of domain. The funnily named log files confused me further and I missed to notice right log file. All of a sudden I was under pressure as business was getting impacted. Luckily, my colleagues Venky and Ashesh stepped in and helped me. We found below error in ohs log file:

[2020-08-24T11:48:33.6940+00:00] [OHS] [ERROR:32] [OH99999] [weblogic] [client_id: 132.240.218.73] [host_id: example.com] [host_addr: 10.0.0.1] [pid: 26738] [tid: 140651159054080] [user: oracle] [ecid: 005fPseYZlA1rYIUIqq2 SR0006 Xb000112] [rid: 0] [VirtualHost: main] <005fPseYZlA1rYIUIqq2 SR0006 Xb000112> Error 2 in opening temp request body file ‘/u01/config/my_domain/config/fmwconfig/components/OHS/instances/tmp/wl_tmprd/_wl_proxy/_post_26738_25’, referer https://example.com/myapp/home

We found that the temp dir location specified with WLTempDir directive was missing based on above error message. So, we created temp directories on all nodes and restarted OHS. It resolved our issue.

I tried to identify what caused disappearance of temp directory. I checked all logs and didn’t find any any information pointing to directory deletion. We raised this with our Linux team to find if any script automatically deleting tmp files. We also raised service request with Oracle and Oracle engineer told us that OHS no way deletes any file. I tried to replicate this issue in lower env, by restarting ohs, weblogic, changing config etc. But I was not able to replicate issue. After hours of analysis, I finally resigned to the fact that RCA can’t be found. So, I prepared a script to notify us if the temp directory gets deleted.

Four months after placing monitoring, we received first alert last week. This gave us opportunity to investigate the issue further. We checked what are the changes done. Only thing that we did on the env was restarting complete domain as part of Disaster Recovery drill. We repeated steps and tmp directory was getting deleted every time.

The hypothesis we derived from this is: We created temp directory directly under OHS instances directory. When admin server was restarted, it tried to sync config from its staging directory ($ADOMAIN_HOME/config/fmwconfig/components/OHS) into mdomain OHS instances directory. Admin Server was treating tmp directory as OHS instance name and as this was not part of list of instances in staging config, it was removing it. Changing tmp directory to location other than instances would resolve this problem. Temp was created this way only in prod and created under different directory in lower envs. Time and again it proved that all envs needs to have same configuration.

Comments

Popular posts from this blog

HOW WE REDUCED SOA OSB PROVISIONING FROM 4 DAYS TO 4 HOURS

NOT ABLE TO START RABBITMQ CLUSTER: CANNOT DECLARE A QUEUE ‘~S’ ON NODE ‘~S’: ~255P

SOA SUITE 12.2.1.4 INSTALLATION: GOT EXCEPTION WHEN AUTO CONFIGURING THE SCHEMA COMPONENT(S) WITH DATA OBTAINED FROM SHADOW TABLE