Posts

KONG API INSTALLATION FAILURE DUE TO CASSANDRA CONFIG

We have Kong API running on Kubernetes. It uses Cassandra to store meta data. The setup used in production is clustered one but due to cost consideration we were asked to build a POC environment with single node. We were given two VMs – one to run Kubernetes control plane and other to run Cassandra and Kong worker nodes. As you are aware there are 3 basic steps to install Kong API: 1) Create and run Cassandra pods. 2) Run Kong migrations to create Kong data schemas in Cassandra. 3) Create and run Kong pods. We made necessary changes in values.yaml of helm charts of Cassandra and Kong to reflect new environment. When we ran Helm charts Cassandra DB and Kong migrations got created successfully but Kong pods were stuck in Init state. Kong pod had two containers:  wait-for-db  and  Kong  containers. wait-for-db container had been stuck so pod also stuck in Init state. We looked at this init container using following command: kubectl logs <kong pod> -c wait-for-db T...

SSL OFFLOADING AT OHS LAYER

  As more and more workloads are moving to security has become more important than ever. SSL over http also know as https is used to encrypt data between client and server i.e., to secure data over wire. It became ubiquitous even for internal applications. Gone are those days doing business over plain http. Though plain is less pain (there is no need to configure and manage certificates), as system admin we need to well equipped with https to secure applications we are managing. Encryption and decryption of data can happen at different layers in Oracle EDG topology which includes external LB, OHS and application layer. Each one have advantages and disadvantages. SSL Offloading at Load Balancer (LB) – LB does all the heavy lifting of SSL handshake and encryption and decryption. As LBs are separate hardware devices it has advantage of using LB hardware and save cpu cycles of application server. Configuration wise too it is most simple solution as SSL certificate configuration at LB i...

SOMEONE WAS EATING MY RAM MEMORY

Today we received an alert on memory utilization alert for couple of VMs where our Kubernetes applications are deployed. We first suspected over utilization of memory by services deployed on K8S. We quickly ran OS commands like top, pidstat to find memory used by our applications. Overall memory usage was hardly crossing 30%. So, we suspected something other than applications like Kernel was consuming memory. Luckily, quick google search gave answer to our problem. The issue was caused by failure to unset Virtual Machine memory ballooning. We ran   vmware-toolbox-cmd stat balloon   to confirm it and issue was reported to concerned team. Below are some references which helped us in finding out problem: https://unix.stackexchange.com/questions/259659/high-memory-usage-but-no-process-is-using-it https://communities.vmware.com/t5/Virtual-Machine-Guest-OS-and-VM/Ballooning-memory-stuck-in-quot-is-resetting-quot/td-p/490441

SSL CERTIFICATE MONITORING

  Expiring SSL certificates are one of the reasons for frequent production incidents and outages. One of the recent example is:  Microsoft Teams . In Java/Weblogic/SOA Suite world there are multiple places where certificates are stored: cacerts, jks trust & identity stores, KSS certificates, Load Balancer certificates etc. If your organization has license for monitoring tools like Oralce OEM, we can setup alerts easily. Our current client doesn’t have OEM. So, we have developed some custom scripts. These scripts can be found in   github .  Luckily, sample scripts for url certificate and jks certificate monitoring are available over internet. So, we modified little bit to suit our requirement. We were not able to find script for monitoring certificates stored OPSS/KSS. So, we have developed a wlst based script.

ASYMMETRIC WEBLOGIC STANDBY DR SITE – ADDITIONAL STEPS REQUIRED FOR SWITCHOVER

Usually asymmetric DR topology is (in which PROD and standby sites’ configurations differ) used to save cost of standby site. If number of instances at both sites are same but number of cpus differ, there won’t be any additional steps required during switchover. But if number of instances differ, switchover will fail due to missing configuration like machines and managed server etc. We got a 4 node weblogic cluster with restricted JRF (no daatbase) at production site and two 2 node cluster at standby site. We followed below additional steps to bring up weblogic domain at standby site. If you have SOA suite or webcenter etc the steps may differ. Removed MS3 and MS4 configuration from config.xml, entries between <server> tags. Removed node 3 and node 4 machine configuration from config.xml, Removed MS3 and MS4 from migratable targets section in config.xml Removed OHS3 and OHS4 configuration from system-components section in config.xml, entries between <system-component> tag. ...

HOW TO USE GUI BASED APPLICATIONS IN LINUX ENVIRONMENTS – PART 1

Image
Sometimes we as sysadmins, need to use Graphical User Interface (GUI) in addition most commonly used ssh terminals. The purpose may be to install software like oracle or accessing graphic applications that are installed only on remote machines. Unix based GUI technologies like X11 or VNC are always confusing to me and I used to rely on my colleagues. Recently, I got to install weblogic on Oracle Cloud Infrastructure (OCI). I invoked installer but no GUI window opened. My colleague, Rohit, configured everything on the VM for me. But this time I wanted to understand how these systems work under the cover so that I can troubleshoot these issues myself next time and save my colleagues’ time :-). In first part of the series, I will explain how to use X11 windows system for GUI in Oracle Cloud Infrasture (OCI). The same steps can be followed in any other Linux environment. In   second part , I will try to explain how things work behind the scenes. For accessing X11 GUI remotely we need t...

WEBLOGIC SERVER GOES TO FAILED_NOT_RESTARTABLE STATE

  We have two node weblogic domain with SOA, OSB, ESS and WSM clusters. Server on one of the nodes was going to FAILED state and was not returning to RUNNING state even after multiple restarts. Logs have following messages: Caused by: java.io.IOException: [DeploymentService:290066]Error occurred while downloading files from Administration Server for deployment request "51,325,278,609,827,905". Underlying error is: "[DeploymentService:290065]Deployment service servlet encountered an Exception while handling the deployment datatransfer message for request id "51,325,278,609,827,905" from server "wls_wsm1". Exception is: "files list is empty"." at weblogic.deploy.service.datatransferhandlers.HttpDataTransferHandler.getDataAsStream(HttpDataTransferHandler.java:92) at weblogic.deploy.service.datatransferhandlers.DataHandlerManager$RemoteDataTransferHandler.getDataAsStream(DataHandlerManager.java:175)weblogic.deploy.internal.tar...