Posts

IT IS DNS AGAIN !!! KONG API WAS STOPPED IN ITS TRACKS

Image
We have a two node Kubernetes cluster with one master and one worker node. Kong api is installed as containers on this cluster. Log files in our Kong API server were located in root directory of the host machine. As the number of logs were getting increased, there was a danger of filling up root directory. So, we decided to move it to separate mount keeping directory path same. Our plan was as below: 1. Cordon and drain the node. 2. Stop kubelet and docker. 3. Move logs to separate mount. 4. Start docker and kubelet. 5. Uncordon the node. We followed the plan but Kong pod was stuck in Init state. We have verified Kong access, error logs, Init (wait-for-db container) container logs and kubelet logs (/var/log/messages in our case). There was nothing suspicious written. We also checked status of all kube-system pods and all were running fine. We verified Control Plane logs as well. We tried restarting docker, kubelet and Cassandra pods (Kong’s persistent store). But nothing worked. Sooner...

WEBLOGIC LOG EXTRACTOR AND SUMMARIZER

Image
  As sysadmins, we need to daily deal with logs. It is overwhelming to check huge log files especially for application servers like weblogic which has numerous log files like diagnostic, server log, out log, OHS logs etc. It is time consuming and quickly lead to fatigue if log files are huge like weblogic. I have developed 2 scripts which extract and summarize logs. I have been using these scripts in both production and test environments. These scripts helped in quickly reacting to a situation and saved time. If your organization already has tools like Splunk, ELK, these scripts may not much use. But for those who couldn’t afford, these scripts could help. Please find brief description about each script as below: Log Extractor : This script takes time in mins or time period as input and extracts logs produced between that time period. We can also filter log files based on log severity as shown below screenshot. Script Location:  https://github.com/RameshPoonati/oracleMiddlewar...

WEBLOGIC ADMIN SERVER STUCK DURING STARTUP

  Recently, we had an issue at VM level which forced us to restart application servers. Most of the weblogic applications started without any issue except one application. When we tried to start admin server, it was hung for a long time. I faced similar issues earlier ( Case 1 ,   Case 2 ,   Case 3 ). So, verified if I ran into one these issues. But luck didn’t favor me and got to investigate it on a Friday evening. I tried restarting again (Friday inertia to deep dive and was looking for a short cut). But it didn’t help so I really got to get my hands dirty. I checked logs but that was not useful much as server.out/nohup.out was stuck after writing below line. <Apr 8, 2021 10:15:45,040 PM PDT> <Warning> <Coherence> <BEA-000000> <2021-04-08 22:15:45.040/85.650 Oracle Coherence GE 12.2.1.3.0 <Warning> (thread=[ACTIVE] ExecuteThread: '0' for queue: 'weblogic.kernel.Default (self-tuning)', member=n/a): The cluster name has not been configur...

KONG API INSTALLATION FAILURE DUE TO CASSANDRA CONFIG

We have Kong API running on Kubernetes. It uses Cassandra to store meta data. The setup used in production is clustered one but due to cost consideration we were asked to build a POC environment with single node. We were given two VMs – one to run Kubernetes control plane and other to run Cassandra and Kong worker nodes. As you are aware there are 3 basic steps to install Kong API: 1) Create and run Cassandra pods. 2) Run Kong migrations to create Kong data schemas in Cassandra. 3) Create and run Kong pods. We made necessary changes in values.yaml of helm charts of Cassandra and Kong to reflect new environment. When we ran Helm charts Cassandra DB and Kong migrations got created successfully but Kong pods were stuck in Init state. Kong pod had two containers:  wait-for-db  and  Kong  containers. wait-for-db container had been stuck so pod also stuck in Init state. We looked at this init container using following command: kubectl logs <kong pod> -c wait-for-db T...

SSL OFFLOADING AT OHS LAYER

  As more and more workloads are moving to security has become more important than ever. SSL over http also know as https is used to encrypt data between client and server i.e., to secure data over wire. It became ubiquitous even for internal applications. Gone are those days doing business over plain http. Though plain is less pain (there is no need to configure and manage certificates), as system admin we need to well equipped with https to secure applications we are managing. Encryption and decryption of data can happen at different layers in Oracle EDG topology which includes external LB, OHS and application layer. Each one have advantages and disadvantages. SSL Offloading at Load Balancer (LB) – LB does all the heavy lifting of SSL handshake and encryption and decryption. As LBs are separate hardware devices it has advantage of using LB hardware and save cpu cycles of application server. Configuration wise too it is most simple solution as SSL certificate configuration at LB i...

SOMEONE WAS EATING MY RAM MEMORY

Today we received an alert on memory utilization alert for couple of VMs where our Kubernetes applications are deployed. We first suspected over utilization of memory by services deployed on K8S. We quickly ran OS commands like top, pidstat to find memory used by our applications. Overall memory usage was hardly crossing 30%. So, we suspected something other than applications like Kernel was consuming memory. Luckily, quick google search gave answer to our problem. The issue was caused by failure to unset Virtual Machine memory ballooning. We ran   vmware-toolbox-cmd stat balloon   to confirm it and issue was reported to concerned team. Below are some references which helped us in finding out problem: https://unix.stackexchange.com/questions/259659/high-memory-usage-but-no-process-is-using-it https://communities.vmware.com/t5/Virtual-Machine-Guest-OS-and-VM/Ballooning-memory-stuck-in-quot-is-resetting-quot/td-p/490441

SSL CERTIFICATE MONITORING

  Expiring SSL certificates are one of the reasons for frequent production incidents and outages. One of the recent example is:  Microsoft Teams . In Java/Weblogic/SOA Suite world there are multiple places where certificates are stored: cacerts, jks trust & identity stores, KSS certificates, Load Balancer certificates etc. If your organization has license for monitoring tools like Oralce OEM, we can setup alerts easily. Our current client doesn’t have OEM. So, we have developed some custom scripts. These scripts can be found in   github .  Luckily, sample scripts for url certificate and jks certificate monitoring are available over internet. So, we modified little bit to suit our requirement. We were not able to find script for monitoring certificates stored OPSS/KSS. So, we have developed a wlst based script.