Posts

HOW WE REDUCED SOA OSB PROVISIONING FROM 4 DAYS TO 4 HOURS

 In our organization, we’ve undertaken the task of upgrading both SOA and OSB from version 12.2.1.3 to 12.2.1.4. This involves an out-of-place migration approach, wherein we create a new environment and switch traffic during the cutover. As you know, this process typically requires several days we opt for manual provisioing. A very significant amount of automation work was previously accomplished for 12.2.1.3 by our predecessors. Despite these efforts, the process still used to take lot of time due to following factors: Bugs in existing automation. For example, SOA composite deployments is not stable which warrants lot of manual work Import of 10s of TLS Certificates into KSS Manual retiring and shutting down of services Monitoring setup Miscellaneous: Increasing Log file count, enabling GC logging, changing heap parameters, bash profile changes While our current automation initiatives have yielded substantial savings, we aimed to capitalize on this upgrade activity to enable end-t...

SOA SUITE 12.2.1.4 INSTALLATION: GOT EXCEPTION WHEN AUTO CONFIGURING THE SCHEMA COMPONENT(S) WITH DATA OBTAINED FROM SHADOW TABLE

  We planned to do out of place upgrade from soa suite 12.2.1.3 to 12.2.1.4. As part of this, we planned to automate end to end provisioning. While executing domain configuration using ansible and python scripts, we got following error in domain creation step. Got exception when auto configuring the schema component(s) with data obtained from shadow table. Failed to build JDBC Connection object: com.oracle.cie.domain.script.jython.CommandExceptionHandler. exception. Quick internet search gave few solutions like deleting domain-registry.xml file, verifying dehydration store password supplied to script etc. But none of them helped. After trying several options, I thought I might have done some stupid error prior to domain creation step. To track down the issue, though of checking what getDatabaseDefaults actually doing. I tried to grep for the getDatabaseDefaults function call inside oracle home. But couldn’t find the source code for the python function. As a final resort, thought of...

RABBITMQ CONNECTION ERROR: JAVAX.NET.SSL.SSLHANDSHAKEEXCEPTION: INVALID ECDH SERVERKEYEXCHANGE SIGNATURE

We have a 7 node test rabbitmq cluster. One of the rabbitmq users reported that they were getting “ javax.net.ssl.SSLHandshakeException: Invalid ECDH ServerKeyExchange signature ” error while connecting to Rabbitmq. We checked logs and also checked application up time. Rabbitmq servers hadn’t been restarted for the past 25 days. Also, we didn’t observe any errors in logs related to vhost used by client. We also checked if the certificate is expired and it was not. We got reasonably confident that it was not our issue and asked to check from client side. As day progressed, we got complaints from two other client application teams. One of them uses php and other uses java. At this point, we started suspecting rabbitmq. Also, one team with multiple consumers told us that they are facing issue only with few of the consumers. With this information, we felt that issue could be with few of the nodes and not complete cluster. One of the blogs on the internet suggested that this error could occ...

NOT ABLE TO START RABBITMQ CLUSTER: CANNOT DECLARE A QUEUE ‘~S’ ON NODE ‘~S’: ~255P

We have 7 node rabbitmq cluster and we are using sharding queues. We have shutdown all non prod rabbitmq nodes as part of VM patching. As it was non prod, we thought of shutting down all together instead of rolling patching. We couldn’t start the any of the rabbitmq post patching. All of nodes were failing to start with the below error: 2022-06-23 05:13:10.104845-07:00 [error] <0.647.0> “Cannot declare a queue ‘~s’ on node ‘~s’: ~255p”, 2022-06-23 05:13:10.104845-07:00 [error] <0.647.0> [“queue ‘sharding: shard.my.queue – rabbit@rabbit1’ in vhost ‘my_vhost'” This is a known issue in some old versions of rabbitmq. rabbitmq node fails to come up with above error. Work around for this is to disable sharding policy for the problematic sharding queue. It can be done either from command line or from rabbitmq admin page. It requires to connect a running node. But as all the nodes were down, we could not disable sharding policy. Luckily, rabbitmq allows disabling sharding pulgi...

TELEGRAF AGENT NOT ABLE TO MONITOR ZOOKEEPER

We use VMWare Wavefront for monitoring and visualization. We use telegraf agent as metrics collector. As part of Kafka monitoring, we have created alerts for Zookeeper availability. It was working fine initially but stopped working after a Kafka upgrade. We checked telegraf agent and zookeeper logs but could not find anything suspicious. We checked telegraf zookeeper github repository. It mentioned that it uses zookeeper mntr command to get monitoring data. We hadn’t whitelisted any zookeeper 4lw commands in earlier versions of Kafka too. It turned out that from Zookeeper version 3.5.3 onwards, we need to explicitly whilelist commands. As mntr is disabled by default, telegraf was not able to collect metrics. Issue got resolved after whitelisting mntr command in zookeeper.properties file and restart of zookeeper.

MONITORING WEBLOGIC, SOA, OSB USING PROMETHEUS AND GRAFANA

Image
  In the   previous blog post , we discussed about how to monitor weblogic based applications including SOA and OSB using Vmware Wavefront. In this blog post, let us explore open source alternative with Prometheus and Grafana. We will use Oracle   Weblogic Monitoring Exporter   in place of jolokia agent to export metric to Prometheus and Grafana for visualizing metrics. High level steps are: 1. Download and install Weblogic Monitoring Exporter. 2. Install Prometheus & Grafana 3. Configure Prometheus to scrape metrics from wls-exporter 4. Configure Grafana dashboards using Prometheus datasource. Below are detailed steps: Install Weblogic Monitoring Exporter : Go to  Weblogic Monitoring Exporter github releases  page and download latest  get_v<version>.sh  script. Copy the exporter configuration file  located here . and pass that as parameter (e.g., ./get_v2.1.2.sh exporter_config.yaml) to get_v<version>.sh. This script download...

WEBLOGIC, SOA AND OSB MONITORING USING WAVEFRONT

Image
  Traditionally sysadmins used bash, WLST/Jython scripts to monitor weblogic based applications including Oracle SOA Suite and Oracle Service Bus. There are few disadvantages with this approach: Need to maintain multiple scripts to monitor a single domain. Sysadmins needs to be proficient in 1 or 2 scripting language. Developing script may take few hours to few days. Will not have access to historical monitoring data for doing trend analysis. With the advent of modern monitoring tools like Prometheus and Grafana, above challenges are addressed. We can setup beautiful graphs and dashboards quickly and tools like Grafana also support alerting mechanism. Wavefront is an integrated solution that can store time series metrics, supports great visualizations, alerting, tracing and more. This blog post describes steps to integrate Weblogic monitoring with Wavefront. Weblogic doesn’t expose time series metrics data directly. So, we need to install Jolokia on Weblogic for exposing JMX metric...

NOT ABLE TO SEE ORACLE SOA COMPOSITE INSTANCES IN EM CONSOLE

Image
We recently built a new Oracle SOA suite environment with version 12.2.1.3. A particular service which is exposed through Gateway was throwing 401 Unauthorized error. There were no instances for this service in em console. So, we initially thought that issue might be with Gateway and analysis was directed towards it. From Gateway logs we found that back end (SOA service) was throwing 401 error. When we checked soa access logs (usually located under server logs directory) we observed 401 errors for this service. But there were no instances in em console. When we tail the logs and hit the service multiple times from postman, we observed below error in logs. <Jan 26, 2023 8:32:05,927 PM PST> <Error> <oracle.wsm.resources.security> <WSM-00008> <Login Exception: [Security:090938]Authentication failure: The specified user failed to log in. javax.security.auth.login.FailedLoginException: [Security:090302]Authentication Failed: User specified user denied.> <Jan...

MONITORING DATASOURCES IN WEBLOGIC/SOA/OSB

Datasource monitoring is one of the key monitoring topics in weblogic based applications. We have been using   simple WLST script   to alert us when a datasource is not working properly. Most of the times, root cause will be a intermittent network/DB issue. To recover from the issue, we will reset/restart the datasource manually. Even though our team is very responsive, sometimes delays in response are inevitable. As one of the tenets of SRE is to automate as much as possible to reduce human toil and manual errors, we have done further automation. We have improved on the earlier version to recover datasource automatically. If the issue can’t be resolved with restart, script will send an email with error message and datasource details. This will save time in gathering error and db information. DBA team, which is part of alert targets, can also quickly act on the alert. New script is available at   this location .

WEBLOGIC MANAGED SERVER START FAILS: “THE NODE MANAGER ASSOCIATED WITH MACHINE SOAHOST1 IS NOT REACHABLE”

We have faced weird issue while starting Weblogic managed server from admin console. We spent good amount of time in analyzing and fixing the issue. Sharing this if it can help others. We made a configuration change in weblogic domain that required to restart all servers including admin server. We did rolling restart of servers. All servers were started except one of the managed servers. While starting the server from console, we got the error: “ the Node Manager associated with machine SOAHOST1 is not reachable “ We logged in into manged server VM and noticed that NM was running fine. Somehow, admin server showing status as unreachable in admin console (Environment -> Machines -> Machine Name -> Monitoring -> Node Manager Status). To confirm that it was not network issue, we tested network connectivity using telnet ( telnet managedServerHost:5556 ). Connection was successful. We checked the Admin server log and found below error: <Nov 16, 2022 11:41:07,357 AM PST> ...

WEBLOGIC ADMIN SERVER FAILING TO START WITH ERROR: SERVICE WEBLOGIC.SERVER.SERVERLIFECYCLESERVICE WAS STARTED AT LEVEL 9 BUT IT HAS A RUN LEVEL OF 10

We had an issue with shared storage server which is used to store weblogic admin server configuration. So, we restarted the admin server once the storage issue was resolved. But restart failed with below errors: <Nov 13, 2022 10:06:15,079 PM PST> <Critical> <WebLogicServer> <BEA-000386> <Server subsystem failed. Reason: A MultiException has 20 exceptions. They are: 1. java.lang.NullPointerException 2. java.lang.IllegalStateException: Unable to perform operation: post construct on weblogic.store.admin.DefaultStoreService 3. java.lang.IllegalArgumentException: While attempting to resolve the dependencies of weblogic.transaction.internal.TransactionService errors were found 4. java.lang.IllegalStateException: Unable to perform operation: resolve on weblogic.transaction.internal.TransactionService 5. java.lang.IllegalArgumentException: While attempting to resolve the dependencies of weblogic.jdbc.common.internal.JDBCService errors were found 20. java.lang.Illeg...

SCG AND VAULT INTEGRATION: LOGIN UNAUTHORIZED DUE TO: X509: CERTIFICATE SIGNED BY UNKNOWN AUTHORITY

We were trying to integrate Spring Cloud Gateway running on K8S with HashiCorp Vault. Wanted to share info how we resolved these issues. Issue 1 : [ERROR] auth.kubernetes.auth_kubernetes: login unauthorized due to: Post “https://10.0.0.:6443/apis/authentication.k8s.io/v1/tokenreviews”: x509: certificate signed by unknown authority Solution : We used following kubernetes auth config to authenticate client to vault: vault write auth/kubernetes/config token_reviewer_jwt=”$SA_JWT_TOKEN” kubernetes_host=”$K8S_HOST” kubernetes_ca_cert=”$SA_CA_CRT” issuer=”https://kubernetes.default.svc.cluster.local” disable_iss_validation=”true” We extracted certificate info using below command: export SA_CA_CRT=$(kubectl config view –raw –minify –flatten –output ‘jsonpath={.clusters[].cluster.certificate-authority-data}’ | base64 –decode) While copying the certificate info to vault container we used  echo $SA_CA_CRT  instead of  echo “$SA_CA_CRT” . Due to this new line characters from certifi...

INTRIGUING 502 BAD GATEWAY ERROR WHILE INVOKING KONG THAT PROXIES BOOMI SERVICE

  Our dev team has exposed a Boomi service through Kong Gateway to be consumed by other teams. The service was working for most of the time but throwing 502 errors randomly. Initially, we thought it was network glitch and ignored those errors. As the testing progressed to higher environments, the error rate increased. Though it was below 1 %, due to limitations at both client and upstream we were tasked to analyze and fix the issue. The high level request flow is like this:   Client -> Kong Gateway -> Boomi VIP -> Boomi Atoms -> Upstream/backend service. We started analysis with Kong and Boomi logs. Found below entries: Kong : {“date”:1.664432446599347E9,”log”:”2022-09-29T06:20:46.599260909Z stderr F 2022/09/29 06:20:46 [error] 38#0: *1693441  upstream prematurely closed connection  while reading response header from upstream, client: 10.200.63.6, server: kong, request: \”POST /test/outbound/consumer HTTP/1.0\”, upstream: Boomi : 2022_09_28.shared_http_se...

TROUBLESHOOTING HASHICORP VAULT KUBERNETES AUTH ERROR

  We were trying to integrate Spring Cloud Gateway (SCG) on Kubernetes with HashiCorp. We followed the steps mentioned in vault   documentation . We were able to bring up vault and vault injector but SCG pods were stuck in init state. Found following error in SCG application pod vault-agent-init container logs: NAME READY STATUS RESTARTS AGE scg-0 0/2 Init:0/1 0 9h vault-0 1/1 Running 0 10h vault-agent-injector-5c89c7dfc5-n2v6v 1/1 Running 0 20h Cgo: disabled Log Level: info Version: Vault v1.8.4 Version Sha: 925bc650ad1d997e84fbb832f302a6bfe0105bbb 2022-09-30T16:24:28.007Z [INFO] sink.server: starting sink server 2022-09-30T16:24:28.007Z INFO creating watcher 2022-09-30T16:25:28.008Z [ERROR] auth.handler: error authenticating: error="context deadline exceeded" backoff=1s To verify that vault injector was connecting to right vault instance, we verified init agent config using below kubectl command: kubectl exec -it scg-0 -c va...

CONNECTION RESET ERROR – BOOMI CONNECTING TO SFTP SERVER

Image
  This blog is about connection reset error we received while connecting from Boomi atom to remote sftp. This issue can occur for any client connecting to a remote sftp server. The sftp server we were trying to connect was old sftp server and it has been used in many integrations. We haven’t faced reset issue for other integrations flow except this new one. So, we first suspected that there was something wrong with the process itself. We looked closely at the process. The process has multiple processes – one main process and multiple child processes. Main process checks for availability of files and child processes transfer files to Boomi atoms for further processing. The first main process and few child processes succeed always but last few process would fail always (usually after processing 7 or 8 files). We tweaked the process to drill down the issue. Here are few thing we tried: Disabled all other processes that connect to the same sftp server. Issue still occurred. We checked ...

HOW TO MONITOR KAFKA CONSUMER LAG

There are couple of ways to monitor Kafka consumer Lag. There is an open source project called   Apache Burrow   which has lot of features and easy to set up. It filters out lot of false positives too. If there are an limitations of using it in your project, we can write bash scripts using Kafka consumer utility which ships with Kafka. We have written one such script. If you are interested,   please check it out here .

KAFKA CONSUMER CLI ERROR: TIMED OUT WAITING FOR A NODE ASSIGNMENT

This post is more of a personal note than a perfect analysis. I have applied lot of workarounds and still not able to identify exactly what caused issue. Writing this blog so it will act as future reference if similar issue happens and add more details about the issue. We were adding monitoring for Kafka consumer lag using kafka-consumer-groups.sh which is located under Kafka installation bin directory. We were able to get lag in all environment except one and script execution failed with below error ( $KAFKA_BIN/kafka-consumer-groups.sh –describe –all-groups –bootstrap-server localhost:9092 ): java.util.concurrent.ExecutionException: org.apache.kafka.common.errors.TimeoutException: Timed out waiting for a node assignment. Call: metadata at java.util.concurrent.CompletableFuture.reportGet(CompletableFuture.java:357) at java.util.concurrent.CompletableFuture.get(CompletableFuture.java:1908) at org.apache.kafka.common.internals.KafkaFutureImpl.get(KafkaFutureImpl.java:165) at kafka.admin...