Posts

MONITORING DATASOURCES IN WEBLOGIC/SOA/OSB

Datasource monitoring is one of the key monitoring topics in weblogic based applications. We have been using   simple WLST script   to alert us when a datasource is not working properly. Most of the times, root cause will be a intermittent network/DB issue. To recover from the issue, we will reset/restart the datasource manually. Even though our team is very responsive, sometimes delays in response are inevitable. As one of the tenets of SRE is to automate as much as possible to reduce human toil and manual errors, we have done further automation. We have improved on the earlier version to recover datasource automatically. If the issue can’t be resolved with restart, script will send an email with error message and datasource details. This will save time in gathering error and db information. DBA team, which is part of alert targets, can also quickly act on the alert. New script is available at   this location .

ISSUES FACED DURING REBUILDING A KUBERNETES CLUSTER

We had a Kubernetes POC cluster with version 1.18. This cluster got corrupted during experimentation by our team. Instead of starting with clean slate, which is comparatively easy, we tried to rebuild cluster with existing kubelet, etcd, kubeadm. This post is basically my own reference/notes of the issues faced and their fixes. Issue 1 :  kubeadm join  failing with below error. This happens only while join control plane nodes not worker nodes. failure loading certificate for CA: couldn’t load the certificate file /etc/kubernetes/pki/ca.crt: open /etc/kubernetes/pki/ca.crt: no such file or directory Solution : Copied following files from node 1 to other control plane nodes /etc/kubernetes/pki/ca.crt /etc/kubernetes/pki/ca.key /etc/kubernetes/pki/sa.key /etc/kubernetes/pki/sa.pub /etc/kubernetes/pki/front-proxy-ca.crt /etc/kubernetes/pki/front-proxy-ca.key Issue 2 : Flannel pods are crashlooping. Found following error continuously in kube-proxy logs.node.go:125] Failed to retrie...

KAFKA BROKER REFUSING TO START – TIMED OUT WAITING FOR CONNECTION

  As part of regular maintenance, we shutdown Kafka. When we tried to start it, it failed with below error: [2021-11-15 05:01:03,824] INFO Socket connection established to mykafka-a1.abc.com/10.0.95.100:2181, initiating session (org.apache.zookeeper.ClientCnxn) [2021-11-15 05:01:03,824] INFO Unable to read additional data from server sessionid 0x0, likely server has closed socket, closing socket connection and attempting reconnect (org.apache.zookeeper.ClientCnxn) [2021-11-15 05:01:03,889] INFO [ZooKeeperClient Kafka server] Closing. (kafka.zookeeper.ZooKeeperClient) [2021-11-15 05:01:03,927] INFO Session: 0x0 closed (org.apache.zookeeper.ZooKeeper) [2021-11-15 05:01:03,927] INFO EventThread shut down for session: 0x0 (org.apache.zookeeper.ClientCnxn) [2021-11-15 05:01:03,930] INFO [ZooKeeperClient Kafka server] Closed. (kafka.zookeeper.ZooKeeperClient) [2021-11-15 05:01:03,933] ERROR Fatal error during KafkaServer startup. Prepare to shutdown (kafka.server.KafkaServer) kafka.zooke...

USEFUL BASHRC FILES

We as sysadmin spend lot of time at command line and most of the time we use few regular commands like going to logs directory, searching history. I have prepared bash aliases that can be used to save few keystrokes for various applications like weblogic, kubernetes, Boomi etc. And these can be downloaded from   github . Please feel free to use and add your own commands/scripts.

UNABLE TO START BOOMI MOLECULE – CANNOT INHERIT FROM FINAL CLASS

  We have a 3 node Boomi cluster that we were not able to start. When we start it and check status status is shown as ‘ Atom is Running’ but after some time atom stops. When we checked container logs we found below error: Nov 11, 2021 4:15:48 AM PST FINE [com.boomi.container.plugin.BasePluginManager createPluginClassLoader] Creating QUEUE_SERVER Plugin ClassLoader with overrides: [] Nov 11, 2021 4:15:48 AM PST SEVERE [com.boomi.container.core.BaseContainer start] Atom startup failed, shutting down Nov 11, 2021 4:15:48 AM PST INFO [com.boomi.util.LogUtil doLog] Atom stop requested. Current status is INITIALIZING Nov 11, 2021 4:15:48 AM PST INFO [com.boomi.container.config.ContainerConfig setStatus] Container status changed from INITIALIZING to STOPPING: Atom startup failed, shutting down Nov 11, 2021 4:15:48 AM PST FINE [com.boomi.schedule.Scheduler removeScheduleGroup] Removing schedule for groupId: internal-dead-attachment-group Nov 11, 2021 4:15:48 AM PST INFO [com.boomi.containe...

ABORTING ORACLE SOA INSTANCES PROGRAMMATICALLY USING WLST

There are multiple ways to abort recoverable instances in SOA suites: From em console Updating dehydration tables Programmatically First option, em console, is useful when there are few instances to be aborted. Though we can update dehydration tables to change instance status, it is not a Oracle recommended approach. I found main tables that are updated when will abort instances by tracing using JMC. But this list is not exhaustive. I am providing sql statements only for reference. Don’t execute these statement unless you understand implications. Aborting instances through SOA management API is a cleaner and quick approach. Please refer to this  git repo  for the script based approach. DELETE FROM MEDIATOR_CALLBACK WHERE FLOW_ID = 27594227; DELETE FROM MEDIATOR_CORRELATION WHERE FLOW_ID = 27594227; UPDATE CUBE_INSTANCE SET SCOPE_REVISION = SCOPE_REVISION + 1 WHERE FLOW_ID = 27594227; UPDATE SCA_COMMON_FAULT SET STATE=256 WHERE FLOW_ID = 27594227 AND STATE <> 2816 AND ST...

IT IS DNS AGAIN !!! KONG API WAS STOPPED IN ITS TRACKS

Image
We have a two node Kubernetes cluster with one master and one worker node. Kong api is installed as containers on this cluster. Log files in our Kong API server were located in root directory of the host machine. As the number of logs were getting increased, there was a danger of filling up root directory. So, we decided to move it to separate mount keeping directory path same. Our plan was as below: 1. Cordon and drain the node. 2. Stop kubelet and docker. 3. Move logs to separate mount. 4. Start docker and kubelet. 5. Uncordon the node. We followed the plan but Kong pod was stuck in Init state. We have verified Kong access, error logs, Init (wait-for-db container) container logs and kubelet logs (/var/log/messages in our case). There was nothing suspicious written. We also checked status of all kube-system pods and all were running fine. We verified Control Plane logs as well. We tried restarting docker, kubelet and Cassandra pods (Kong’s persistent store). But nothing worked. Sooner...