Posts

Showing posts with the label Oracle SOA

Server state changed to FAILED - Error occurred while downloading files from Administration Server for deployment request

Oracle weblogic managed servers (SOA/OSB/ESS) were failing to restart after making few changes. The error written into logs was not straight forward and we spent couple of hours debugging the issue. Sharing this so that others can benefit from it. We are using Exadata as dehydration store. Network domain for applications we are managing got changed. As part of this change we changed datasource, ftp adapter config to refer to new domain. We made all required changes and tried to restart servers. We were able to bring up admin server but all managed servers failed to restart. Following error was written logs. Caused by: java.io.IOException: [DeploymentService:290066]Error occurred while downloading files from Administration Server for deployment request “42,336,427,916,066,804”. Underlying error is: “[DeploymentService:290065]Deployment service servlet encountered an Exception while handling the deployment datatransfer message for request id “42,336,427,916,066,804...

SSL/TLS Certificates and Weblogic

Let us learn about SSL/TLS and how it applies to Weblogic and Fusion Middleware. 1) What is SSL/TLS? Transport Layer Secuity (TLS) and Secure Socket Layer (SSL) are cryptographic protocols used to securely communicate over a computer network. SSL was deprecated in 2014 after finding a vulnerability. TLS became new standard for secured http. TLS 1.3 is latest TLS version. 2) Why SSL/TLS? It is easy to intercept and read data transmitted in plain text over a network. Obviously we don’t like bank details, email or personal information fall in wrong hands. TLS protects from this by encrypting data between client and server.   3) How TLS works? TLS uses public key infrastructure to encrypt data exchange. TLS handshake at high level:      i)    Agree on version of TLS to use.      ii)   Agree on Cipher suites to use.      iii) Validate identity of the server using server certificate by client. In case of two way ...

Struck with Stuck Threads Running Across EBS, SOA and ODI

There was an interesting issue that involved EBS, SOA and ODI with a client that I worked earlier. Customer is using Product Data Hub (PDH) application along with Integrated SOA Gateway (ISG) for publishing product details from master repository (EBS) to downstream systems.  Oracle SOA is used as integration layer and ODI is used to extract bulk data from PDH. All applications are deployed in Azure and Azure Application Gateway is used as Loadbalancer (LB). High level flow is as below: EBS Business Event -> SOA -> ODI PIM Webservice -> ODI Scenario -> EBS DB. Lately the customer has been observing stuck threads in SOA and ODI and response never comes back to SOA even after waiting for hours. I got interested in this problem, will let you know the reason at the end. We followed below procedure to find the root cause: The issue was replicated at will in lower envs. We observed stuck threads in both SOA and ODI when issue occurred. We took the thread dumps. We could see th...

Let us Patch Up with Patching - weblogic/soa suite

Patching is a confusing topic for beginners. In this post, I will try to explain different terminologies. CPU – Short for Critical Patch Update. It comprises of security patches.   PSU – Short for Patch Set Updates. It includes security and priority fixes. Both patches are released quarterly on the Tuesday closest to the 17th day of the months of January, April, July and October. Both the patches are cumulative that means every patch includes fixes of the early patches. So, there is no need to apply earlier patches. Once a PSU has been applied on the system, the recommended method to apply all future CPU program security content is to apply future PSUs. In other words, once a PSU has been applied, it is not recommended to switch back to traditional N-apply patches. Reverting to CPU from PSU is complex and time consuming task and so it is not recommended. One-Off : If Oracle finds critical vulnerability that can’t wait till next patch release, it rel...

Useful SQLs for Querying SOA Suite Dehydration Tables

As you all know Oracle SOA Suite uses dehydration tables to store the state of instances. We can query these tables to troubleshoot issues. Here are few sqls I find useful: 1) Count of bpel instances created between two timestamps. Count doesn’t include mediator instances: select count(1) from cube_instance  where creation_date between to_timestamp('2018-04-23 09:00', 'YYYY-MM-DD HH24:MI') and to_timestamp('2018-04-24 17:00', 'YYYY-MM-DD HH24:MI') 2) Count of bpel instances created between two timestamps group by hour: select to_char(creation_date, 'YYYY-MM-DD HH24'), count(1) from cube_instance where creation_date between to_timestamp('2018-04-23 09:00', 'YYYY-MM-DD HH24:MI') and to_timestamp('2018-04-24 17:00', 'YYYY-MM-DD HH24:MI') group by to_char(creation_date, 'YYYY-MM-DD HH24') order by to_char(creation_date, 'YYYY-MM-DD HH24'); 3) Average and max response times: select TO_CHAR(created_time, ...

Useful Unix Commands and Tips - Part 4

Let us learn about tar and ps in this part of the Unix series: tar : tar is derived from Tape Archive when magnetic tapes were used to archive files. tar utility is useful in creating single archive file from multiple files maintaining owners, file types, permissions etc. It is mainly used to create backups and copy files from one computer to other. Let us see few examples: 1) To create tar file from a dir: tar -cvf <tar file name> <Source dir name> #Here, c stands for compress, v for verbose, f - file name of tar file. tar -cvf scripts.tar /opt/oracle/product/admin/scripts 2) To extract files from tar file: tar xvf scripts.tar 3) tar doesn’t compress files, it only creates archive file. It can be compressed using gzip utility. gzip scripts.tar # It creates a compressed file scripts.tar.gz gzip -d scripts.tar.gz  # to unzip 4) We can archive and compress with single command using ‘z’ option: tar -zcvf scripts.tar.gz scripts # Compress tar -zxvf scripts...

Useful Unix Commands and Tips - Part 2

This is in continuation to the first  part in the series. Let us discuss grep and lsof in part 2: grep : 1) Case insensitive search: grep -i "string" FILE   2) Checking for full words, not for sub-strings using grep -w grep -w "serachword" file 3) Displaying lines before/after/around the match using grep -A, -B and -C. This is very useful if you want to see lines around search string: i) grep -A <N> "string" FILENAME # display N lines after match: ii) grep -B <N> "string" FILENAME # display N lines before match iii) grep -C 2 "Example" demo_text # display N lines around match 4) Searching in all files recursively using grep -r grep -r "ramesh" * 5) Invert match using grep -v grep -v -e "pattern" -e "pattern" 6) Counting the number of matches using grep -c. There is no need to pipe the output to wc. grep -c "pattern" filename grep -v -c this demo_file (Count of lines not matching patter...

strace to debug slowness/hang issues

strace is a Unix diagnostic tool used to trace system calls like read, write, connect etc. It comes handy for analyzing system hang and slowness issues. One of the key advantages is that we can use this tool without root access. But it comes with disadvantage too, it greatly slows down the process we are tracing. So don’t use it in production systems.    strace generates lot of output, at the beginning it will be quiet overwhelming. As you work on problems, you will get comfortable using the tool. Let us see some examples of strace.     1) To check what system calls a command is making:    strace <command> e.g., strace ls   2) If the process is already started and want to debug diagnose problems:    strace -p <pid>   3) We can save the output of strace to a file using -o option e.g., strace -o out.txt ls   4) To trace a bash script: strace -f -o trace.out bash <script name>   5) To trace specific system calls instead tracing all systems calls, we can use -e optio...

Its time for Logging - Log file locations, Logger Levels for Weblogic, SOA and other middleware products

The first port of call for analyzing most of the issues is log files. Lot of beginners try to resolve an issue by looking at error message generated in application, console etc.That may not give you exact error always. So, it is better to look at log files as they provide lot more information to solve the problem. In this blog, let us look at different log files available in Weblogic/Middleware and how increase log levels. It is good practice is note down all log file locations of systems we are supporting and ways to increase log levels. Diagnostic Logs : Logs are written using ODL format. Most of the SOA, OSB and other application  built on top of weblogic are written to this log file. So, if you are working on any of those products, it is better to look at this log file first. The default log file location is: $DOMAIN_HOME/servers/$SERVER_NAME/logs. If the information written in the logs is not sufficient to anlayze the issue, we can increase logger levels from em console. We can al...

wlst and shell scripts for monitoring weblogic, soa suite, oracle service bus and others

Scripts to monitor and managed Oracle Fusion Middleware applications are available at following github location:  https://github.com/RameshPoonati/oracleMiddlewareScripts Weblogic : Disable Diagnostic Modules Error Report Generation from logs Application/Deployment Monitoring http log monitoring JMS Bridge Monitoring JMS Consumer Monitoring Log Archiving JMS Queue Monitoring Server Monitoring Domain Startup/Shutdown Thread Dump Generation Thread Monitoring SOA Suite : Monitor Adapter Status Reference instance response time from audit trail ODI : Duplicate Execution Monitoring Oracle Service Bus (OSB) : OSB Proxy Status Change Shell Scripts : Disk Usage Monitoring

How to take Java Thread and Heap Dumps

Weblogic is nothing but a java process. So like any other java process we can take Thread and Heap dumps to analyze issues. Let us look at them one by one. Thread Dump : Thread dump gives a snapshot of threads at the time of taking dumps. It is primarily used to analyze performance and lock/resource contention issues. It is recommended to take 5-6 thread dumps in intervals of 10-15 seconds to understand if there is any issue. With single thread dump we won’t be able judge if it contention is momentarily or of long duration. There are multiple ways of taking thread dumps but my personal choice is jstack as we can take dumps even during most of the jvm hangs. 1) Using WLST : This method is not recommended as weblogic doesn’t respond during high loads. Also, wlst uses lot of resources. Following code snippet is taken from Oracle documentation: Place following code in a file code thread_dump.py. Please make changes according to your setup. serverName = ‘soa_server1’ ...

Time for Timeout !!!

One of the intriguing topics in Oracle SOA Suite is timeout. Most of the content here is taken from Oracle and other blogs. Please check references for more information. 1) syncMaxWaitTime : As per Oracle Doc ID 880313.1: The SyncMaxWaitTime property applies to durable processes that are called in an asynchronous manner. When the client (or another BPEL process) calls the BPEL process which has durable activity(ex : wait), the wait (breakpoint) activity is executed. However, since the wait is processed after some time by an asynchronous thread in the background, the executing thread returns to the client side. The client (actually the delivery service) tries to pick up the reply message, but it is not there since the reply activity in the process has not yet executed. Therefore, the client thread waits for the SyncMaxWaitTime seconds value. If this time is exceeded, then the client thread returns to the caller with a timeout exception.If the wait is less than the SyncMaxWaitTime va...

502 Bad Gateway Error when Oracle WebCenter Lift and Shifted to Azure Cloud

Our client had decided to move most of the OnPrem applications to Azure Cloud to save costs. As part of this we were asked to migrate multiple Oracle Middleware applications to Azure Cloud. One of them was Oracle WebCenter. Though the issue reported here is applicable to any Oracle Middleware products. Our setup has LB, two Oracle http servers and 2 WebCenter servers. We have cloned DB, OHS and WebCenter applications to Azure cloud. After few changes, we were able to bring up admin console and managed servers but we were not able to access console using LB, we were getting 502 Bad Gateway error from Azure Application Gateway. But same urls were perfectly working fine when we either hit ohs or admin server ips directly. The issue was dragged for couple of days and we were worried a lot as we had very tight timelines. We have analyzed server logs but didn’t find anything suspicious except that few errors in OHS access logs. But these http requests were not sent by us as part of our testi...

Not able to abort or complete instances from em console

In one of our test environments, SOA instances were in recovery mode and not able to abort or mark it as complete from em console. em console was timing out if we click on instance.  Root Cause: Thousands of faults occurred for these instances and SOA is keep on retrying. Updating following tables resolved issue. update cube_instance set state = ‘5’  where  cikey in (‘23923268’); update sca_flow_instance set active_component_instances = ‘0’  where  active_component_instances <> ‘0’;

JNDI name is null or empty error while deploying DBAdapter in Oracle SOA Suite

We updated DB Adapter outbound connection pool configuration and tried to deploy it and got below error: <20-Apr-2018 11:18:21 o’clock BST> <Failed to initialize the application “DbAdapter” due to error weblogic.application.ModuleException: weblogic.connector.exception.RAException: JNDI name is null or empty. weblogic.application.ModuleException: weblogic.connector.exception.RAException: JNDI name is null or empty.         at weblogic.application.internal.ExtensibleModuleWrapper.prepare(ExtensibleModuleWrapper.java:114)         at weblogic.application.internal.flow.ModuleListenerInvoker.prepare(ModuleListenerInvoker.java:100)  Root Cause: Save outbound connection pool without XA datasource JNDI name and deployed with new plan. Resolution: Used below command to find entries without JNDI name and deleted thsoe entries. grep -A 1 “ JNDIName ” DB_Plan.xml | more

When DB Indexes were not rebuilt - Unable to get a connection for pool

We had DB based partition for dehydration store and we used to drop old partitions every month to purge old data. We dropped partition as usual but forgot to rebuild indexes. Due to this indexes became unusable and connection pool size of db adapter reached maximum value (1000 in this case). Below error was written to log files.  [2017-07-01T05:49:45.884-07:00] [osb_server1] [ERROR] [OSB-381967] [oracle.osb.transports.jca.jcatransport] [tid: [ACTIVE].ExecuteThread: ’96’ for queue: ‘weblogic.kernel.Default (self-tuning)’] [userId: ] [ecid: 79e678b0-d24f-4750-938a-8fb9e6c2fab2-01e97c36,1:5426731] [APP: Service Bus JCA Transport Provider] [FlowId: 0000LnyCmqJ6YNBrVWYBU81PAaVR00fxSh] Invoke JCA outbound service failed with application error, exception: com.bea.wli.sb.transports.jca.JCATransportException: oracle.tip.adapter.sa.api.JCABindingException: oracle.tip.adapter.sa.impl.fw.ext.org.collaxa.thirdparty.apache.wsif.WSIFException: servicebus:/CommonServices/Ar...

TLS Certificate that lead to team wars

Image
This happened some time around Apr’2019. This story has all features of a suspense movie with a chaotic first half, surprise intermission/interval and fast paced unraveling second half.     Our client has Oracle EBS/ERP and it is integrated with other applications using Oracle middleware (java based) in Azure cloud. ERP admins clone ERP test environment from prod once in a month. Post cloning they do bunch of config updates to test environment. They also do a service re-deployment in our middleware application from ERP. This service is like a gateway to ERP application.   Architecture diagram is as shown below. ERP application has two nodes and is behind loadbalancer (LB). SSL is offloaded at LB. While servicing requests, middleware contacts ERP application via LB. During deployment, ERP send executable jar file to middleware, middleware does lot of validations including validation of ERP endpoint which is exposed through LB (https). If all validations are passed, middlew...

MDB application soa-infra is NOT connected to messaging system

One of our soa servers health went into WARNING state. Also, in application health tab of admin console, we could see ‘MDB application soa-infra is NOT connected to messaging system’ error message. While analyzing server logs, found below error messages: <27-Jun-2019 15:01:03 o’clock BST> <[ACTIVE] ExecuteThread: ’54’ for queue: ‘weblogic.kernel.Default (self-tuning)’> <> <> <23aac39c-aac6-4725-9f14-6f7fd98a1a7c-000001a5> <1561644063760> <[severity-value: 16] [rid: 0] [partition-id: 0] [partition-name: DOMAIN] > <JMS Module “SISJMSModule” deployed to WebLogic domain sod_domain defines an entity of type Uniform Distributed Topic with name “SISJMSModule!CSFNotificationTopic” with default targeting enabled. There is no JMS server having a persistent store with distribution policy “Distributed” available in WebLogic domain sod_domain to host the entity. The destination “SISJMSModule!CSFNotificationTopic” will be hosted when a JMS server having a...

Linux Core files are not generated even after changing ulimit

One of our Oracle SOA servers were crashing lately and Oracle requested for core dumps to analyze issue further. In hotspot crash report, following message is written: Not able to take core dump as core file size is set to zero. Below are steps followed to generate core dumps on JVM crash: Step 1: Increased ulimit using following command: ulimit -c unlimited > /dev/null 2>&1 Step 2: Restarted server. But no luck, still no core dumps are generated. It was not effective as it is applicable to only for the current shell session and not to new shell session. Step 3: Updated the value in .bashrc file ‘ulimit -c unlimited > /dev/null 2>&1′ . This made all new sessions’ ulimit to new value. Again started soa server, still ulimit value for core dump file size is zero. Step 4: This was not expected. On analysis found that, soa server is taking its ulimit values from its parent process which is node manager process as server was started via node manager. Restarted node manage...

BINDING.JCA-12141 error while connecting sonicmq

We have Oracle SOA 11g that connects to sonicmq using JMS adapter. This integration stopped working and found following BINDING.JCA-12141 error in server logs. Followed below procedure to troubleshoot the issue: 1) Checked if the connection factory details are correct or not. 2) Tested connection from admin console sonic queue connection factory of jms adapter. It is also succeeded. 3) Went to Monitoring -> Outbound Connection Pools of Jms Adapter. There checked ‘Rejected Connections’ of eis/sonicmq/Queue. It has high values of rejected connections. 4)Increased log levels of adapters fro em console. Changed log levels of soa.adapter to lowest level. Then observed time out error in diagnostic log. 5) Reported issue to team who is taking care of sonicmq. They found out that there was as an issue with their loadbalancer used in aia connection factory.