Posts

Showing posts with the label Oracle Middleware

Server state changed to FAILED - Error occurred while downloading files from Administration Server for deployment request

Oracle weblogic managed servers (SOA/OSB/ESS) were failing to restart after making few changes. The error written into logs was not straight forward and we spent couple of hours debugging the issue. Sharing this so that others can benefit from it. We are using Exadata as dehydration store. Network domain for applications we are managing got changed. As part of this change we changed datasource, ftp adapter config to refer to new domain. We made all required changes and tried to restart servers. We were able to bring up admin server but all managed servers failed to restart. Following error was written logs. Caused by: java.io.IOException: [DeploymentService:290066]Error occurred while downloading files from Administration Server for deployment request “42,336,427,916,066,804”. Underlying error is: “[DeploymentService:290065]Deployment service servlet encountered an Exception while handling the deployment datatransfer message for request id “42,336,427,916,066,804...

SSL/TLS Certificates and Weblogic

Let us learn about SSL/TLS and how it applies to Weblogic and Fusion Middleware. 1) What is SSL/TLS? Transport Layer Secuity (TLS) and Secure Socket Layer (SSL) are cryptographic protocols used to securely communicate over a computer network. SSL was deprecated in 2014 after finding a vulnerability. TLS became new standard for secured http. TLS 1.3 is latest TLS version. 2) Why SSL/TLS? It is easy to intercept and read data transmitted in plain text over a network. Obviously we don’t like bank details, email or personal information fall in wrong hands. TLS protects from this by encrypting data between client and server.   3) How TLS works? TLS uses public key infrastructure to encrypt data exchange. TLS handshake at high level:      i)    Agree on version of TLS to use.      ii)   Agree on Cipher suites to use.      iii) Validate identity of the server using server certificate by client. In case of two way ...

Struck with Stuck Threads Running Across EBS, SOA and ODI

There was an interesting issue that involved EBS, SOA and ODI with a client that I worked earlier. Customer is using Product Data Hub (PDH) application along with Integrated SOA Gateway (ISG) for publishing product details from master repository (EBS) to downstream systems.  Oracle SOA is used as integration layer and ODI is used to extract bulk data from PDH. All applications are deployed in Azure and Azure Application Gateway is used as Loadbalancer (LB). High level flow is as below: EBS Business Event -> SOA -> ODI PIM Webservice -> ODI Scenario -> EBS DB. Lately the customer has been observing stuck threads in SOA and ODI and response never comes back to SOA even after waiting for hours. I got interested in this problem, will let you know the reason at the end. We followed below procedure to find the root cause: The issue was replicated at will in lower envs. We observed stuck threads in both SOA and ODI when issue occurred. We took the thread dumps. We could see th...

Lonely Admin Server Refused to Start

We had an interesting problem over weekend and thought of sharing with all. One of my colleagues called me up and told that admin server in OSB cluster was not starting up and process was hanging. Here are the steps we followed to troubleshoot the issue. 1) He told that server had been up more than a week and he added JMX port before restarting the server. He was very confident that it was not causing issue. But I insisted to rollback this recent change and start again. Situation was not improved, server startup was hanging again. So, we rule out that this change was not causing issue. 2) We ran top command and checked if there was any pressure on resources. cpu, memory and load average were very low. 3) Checked log files for potential errors. Nothing useful information was found in logs too. 4) The symptoms were similar to recent issue I have worked on. So, we ran lsof command to find if server was waiting for any network connection ( lsof -a -i4 -i6 -itcp -p <pid> ).  lsof...

Server log files mysteriously disappeared !!!

Recently I started working on a new client engagement. This setup was new to us. I was troubleshooting some issue and thought of checking log files. I was able to find server.log and diagnostic log files but surprisingly server.out files were missing all together. Here are the steps I have followed to troubleshoot the issue: 1) Checked if the managed server was started from console (node manager) or using startup scripts. As you know that stdout logs will be written to nohup if started from commandline. Server was started from node manager, so this possibility was ruled out. 2) Compared with logging config of an env where server.out logs are present. Didn’t find any difference. 3) As it is dev env, thought of checking it quickly restarting server. Restarted server and I could see server.out file and it was getting updated. 4) Returned to my desk after lunch, surprisingly the log was gone again. 5) Checked lsof output for this process and found interesting thing as below: [myid@abc...

Let us Patch Up with Patching - weblogic/soa suite

Patching is a confusing topic for beginners. In this post, I will try to explain different terminologies. CPU – Short for Critical Patch Update. It comprises of security patches.   PSU – Short for Patch Set Updates. It includes security and priority fixes. Both patches are released quarterly on the Tuesday closest to the 17th day of the months of January, April, July and October. Both the patches are cumulative that means every patch includes fixes of the early patches. So, there is no need to apply earlier patches. Once a PSU has been applied on the system, the recommended method to apply all future CPU program security content is to apply future PSUs. In other words, once a PSU has been applied, it is not recommended to switch back to traditional N-apply patches. Reverting to CPU from PSU is complex and time consuming task and so it is not recommended. One-Off : If Oracle finds critical vulnerability that can’t wait till next patch release, it rel...

Useful SQLs for Querying SOA Suite Dehydration Tables

As you all know Oracle SOA Suite uses dehydration tables to store the state of instances. We can query these tables to troubleshoot issues. Here are few sqls I find useful: 1) Count of bpel instances created between two timestamps. Count doesn’t include mediator instances: select count(1) from cube_instance  where creation_date between to_timestamp('2018-04-23 09:00', 'YYYY-MM-DD HH24:MI') and to_timestamp('2018-04-24 17:00', 'YYYY-MM-DD HH24:MI') 2) Count of bpel instances created between two timestamps group by hour: select to_char(creation_date, 'YYYY-MM-DD HH24'), count(1) from cube_instance where creation_date between to_timestamp('2018-04-23 09:00', 'YYYY-MM-DD HH24:MI') and to_timestamp('2018-04-24 17:00', 'YYYY-MM-DD HH24:MI') group by to_char(creation_date, 'YYYY-MM-DD HH24') order by to_char(creation_date, 'YYYY-MM-DD HH24'); 3) Average and max response times: select TO_CHAR(created_time, ...

Useful Unix Commands and Tips - Part 7

This is the final part of the series. Let us see some miscellaneous commands and tips. Hope this series helped you. 1) The utilities zcat, zgrep, zless, and zdiff, among others, serve the same purpose as cat, grep, less, and diff, respectively, but operate on compressed files. For example, if you want to search inside jar/zip file, you can use zgrep. 2)  apropos command can be used to find the appropriate command if we forget command name.  apropos search # gives all related commands with search string. 3) watch command allows to watch the output of a program change over time, which can be especially convenient when looking for changes to file sizes that reflect debugging output, download or file creation state, memory or disk usage, and so on. 4) History: If you want to remove a sensitive command from your history, you can simply edit your $HISTFILE history file and remove it. There is a trick you can use if you want to fly under the radar and never have a command recor...

Useful Unix Commands and Tips - Part 6

Let us learn about tools that can help in troubleshooting performance related problems: 1) top : Top provides statistics about cpu, memory utilization, load averages, resident memory, virtual memory etc. The screen refreshes automatically.  i) top u oracle to see all the processes initiated by user oracle. ii) Press ‘1’ to see the individual cores. iii) top -p <pid> : To monitor only processes with a given process id. iv) top -H -p <pid>: Show all threads.We can press H on top output screen as well to get thread info. v) -b Batch mode. Useful for sending output from top to other programs or to a file. vi) To sort memory usage in ‘top’ output, press ‘SHIFT+m’  vii) List of process are sorted and displayed by default based on CPU utilization. We can change it by pressing “<” or “>”, sort column changes accordingly. top command consumes reasonable cpu. So, don’t run too many top sessions. 2) htop : htop is similar to top...

Useful Unix Commands and Tips - Part 5

Let us learn about some miscellaneous commands in this part: 1) To find Linux version of OS: lsb_release -a 2) To find kernel version: uname -a 3) To find number cpu cores: lscpu cat /proc/cpuinfo 4)  more : cat command can be used to view the file contents. But if the file is more than few pages, reading through the buffer is difficult. The commands more and  less are more useful in these cases. We can go to next page by pressing space bar.  By pressing b we can go back by one page. We can search for a string by typing /, followed by the string similar to vi. We can find the next occurrence of the string by typing n. 5) less : less command has similar functionality as more but it is faster. 6) which   command can be used to find the location of a program. which perlwhich java 7) To read a file backwards, we can use tac (reverse cat) and less together: tac nohup.out | less 8) To find packages installed using rpm : rpm -qa | egrep "java-1.7.0-openjdk|java-1.8.0-openjdk 9) fuser i...

Useful Unix Commands and Tips - Part 4

Let us learn about tar and ps in this part of the Unix series: tar : tar is derived from Tape Archive when magnetic tapes were used to archive files. tar utility is useful in creating single archive file from multiple files maintaining owners, file types, permissions etc. It is mainly used to create backups and copy files from one computer to other. Let us see few examples: 1) To create tar file from a dir: tar -cvf <tar file name> <Source dir name> #Here, c stands for compress, v for verbose, f - file name of tar file. tar -cvf scripts.tar /opt/oracle/product/admin/scripts 2) To extract files from tar file: tar xvf scripts.tar 3) tar doesn’t compress files, it only creates archive file. It can be compressed using gzip utility. gzip scripts.tar # It creates a compressed file scripts.tar.gz gzip -d scripts.tar.gz  # to unzip 4) We can archive and compress with single command using ‘z’ option: tar -zcvf scripts.tar.gz scripts # Compress tar -zxvf scripts...

Useful Unix Commands and Tips - Part 3

Let us learn about df, du, sftp and scp commands in this part of the series: du : 1) To know the percentage of disk space used at mount level:  df -kh 2) To sort disk usage in a directory in human readable format: du -sh * | sort -rh 3) If multiple mounts uses common path. du command on common path includes storage used in all mounts. If we want find disk used by only common directory path and not device mounted on common path, need to use –exclude clause of du as shown above. du -sh --exclude=/directory {target-directory/*} 4) To find largest directories:  du -S | sort -n | tail -20 5) To find directories with size greater than 1 GB: du -h --max-depth 50 /orasoa/common/ 2>/dev/null | grep '[0-9]G>' | sort -h 6) du usually don’t include hidden directories/files during disk usgage calculation. To include hidden files/dirs: du -sk * .[a-zA-Z0-9]* 7) -d option prevents du from following file system boundaries. For example, to determine disk usage on the root fil...

Useful Unix Commands and Tips - Part 2

This is in continuation to the first  part in the series. Let us discuss grep and lsof in part 2: grep : 1) Case insensitive search: grep -i "string" FILE   2) Checking for full words, not for sub-strings using grep -w grep -w "serachword" file 3) Displaying lines before/after/around the match using grep -A, -B and -C. This is very useful if you want to see lines around search string: i) grep -A <N> "string" FILENAME # display N lines after match: ii) grep -B <N> "string" FILENAME # display N lines before match iii) grep -C 2 "Example" demo_text # display N lines around match 4) Searching in all files recursively using grep -r grep -r "ramesh" * 5) Invert match using grep -v grep -v -e "pattern" -e "pattern" 6) Counting the number of matches using grep -c. There is no need to pipe the output to wc. grep -c "pattern" filename grep -v -c this demo_file (Count of lines not matching patter...

Useful Unix Commands and Tips - Part 1

In this series of blog posts, I will provide useful Unix commands and tips that are useful in everyday sysadmin life in general and weblogic/oracle middleware admins in particular. I am writing these blogs based on personal notes gathered over years mostly from web. Unfortunately, I didn’t maintain source of the articles/websites. Extremely sorry for not giving due credits. Let us start the series with find command. I am dedicating whole blog post to find as it is versatile and useful. find command is used to search files/directories in a directory. Using find exec option e.g., renamme, copy. 1) To find all files and directories in a given directory and sub directories. find . # Finds all the files and directories in current directory.It searches current dir if none is specified find /abc/xyz # Finds all the files and directories in directory named /abc/xyz 2) We can limit search results only to either files or directories. find . -type f # limits the results to only files. f here...

strace to debug slowness/hang issues

strace is a Unix diagnostic tool used to trace system calls like read, write, connect etc. It comes handy for analyzing system hang and slowness issues. One of the key advantages is that we can use this tool without root access. But it comes with disadvantage too, it greatly slows down the process we are tracing. So don’t use it in production systems.    strace generates lot of output, at the beginning it will be quiet overwhelming. As you work on problems, you will get comfortable using the tool. Let us see some examples of strace.     1) To check what system calls a command is making:    strace <command> e.g., strace ls   2) If the process is already started and want to debug diagnose problems:    strace -p <pid>   3) We can save the output of strace to a file using -o option e.g., strace -o out.txt ls   4) To trace a bash script: strace -f -o trace.out bash <script name>   5) To trace specific system calls instead tracing all systems calls, we can use -e optio...

Its time for Logging - Log file locations, Logger Levels for Weblogic, SOA and other middleware products

The first port of call for analyzing most of the issues is log files. Lot of beginners try to resolve an issue by looking at error message generated in application, console etc.That may not give you exact error always. So, it is better to look at log files as they provide lot more information to solve the problem. In this blog, let us look at different log files available in Weblogic/Middleware and how increase log levels. It is good practice is note down all log file locations of systems we are supporting and ways to increase log levels. Diagnostic Logs : Logs are written using ODL format. Most of the SOA, OSB and other application  built on top of weblogic are written to this log file. So, if you are working on any of those products, it is better to look at this log file first. The default log file location is: $DOMAIN_HOME/servers/$SERVER_NAME/logs. If the information written in the logs is not sufficient to anlayze the issue, we can increase logger levels from em console. We can al...

Oracle MFT Error - Error occurred while polling the listening source

We can setup email notifications whenever an error occurs in MFT like failed connection to source or target etc. This is very useful feature of MFT. We had setup this monitoring for our MFT servers. We used to get below error frequently in production server while connecting to a 3rd party application.   Error occurred while polling the listening source XYZ Cause: Unexpected error occurred while polling the JCA source.                  Action: Review the diagnostic log for the exact error, which begins with the string 'ERROR:'. This error line, along with the preceding lines, indicates the reason for this exception. Contact Oracle Support Services if the problem persists. The endpoint here was the target system to which MFT transfers file. We had setup retry at target so transfers were retried whenever this error occurs. So, it didn’t have direct impact on business. But we used to get around 100 alerts mails with these connection errors. It was tiring experience to read and chec...

wlst and shell scripts for monitoring weblogic, soa suite, oracle service bus and others

Scripts to monitor and managed Oracle Fusion Middleware applications are available at following github location:  https://github.com/RameshPoonati/oracleMiddlewareScripts Weblogic : Disable Diagnostic Modules Error Report Generation from logs Application/Deployment Monitoring http log monitoring JMS Bridge Monitoring JMS Consumer Monitoring Log Archiving JMS Queue Monitoring Server Monitoring Domain Startup/Shutdown Thread Dump Generation Thread Monitoring SOA Suite : Monitor Adapter Status Reference instance response time from audit trail ODI : Duplicate Execution Monitoring Oracle Service Bus (OSB) : OSB Proxy Status Change Shell Scripts : Disk Usage Monitoring

Disk space is not freed in Weblogic even after deleting huge number of files

We had a SOA service which uses mail server to send notifications. One day, mail server went down and this triggered huge amount of log messages in server.out log file. Soon we got monitoring alert saying disk space is 90%. Out log files more than 1 GB size got created. We have had cleared lot of server.out log files. Surprisingly, it didn’t reclaim any space and disk utilization was keep on climbing.  We ran lsof command and grep for server.out file. [oracle@abcd logs]$ lsof -p 6913 | grep soa_server1.out lsof: WARNING: can't stat() tracefs file system /sys/kernel/debug/tracing       Output information may be incomplete. java    6913 oracle    1w      REG              249,0 655131937   3283212 /u01/data/domains/soa_domain/servers/soa_server1/logs/soa_server1.out (deleted) java    6913 oracle    2w      REG            ...

Not Able to Start Weblogic/MFT Server After Out of Memory Error

There is a known bug in Oracle Managed File Transfer (MFT) which causes leaking of connections. This used to cause OOM error once in a month. We have an alert – which sends mail if heap usage is greater than 95% and 3 consecutive garbage collections not able to free much heap space. As we have two node cluster, we restart the server which reached heap utilization of more than 95%. On that particular day, the person on shift missed the alert. He checked alter after sometime and server health became ‘Not Reachable’. He tried to restart from console but he was not able to stop it and server status was showing as ‘RUNNING’. He tried to restart from command line even it didn’t work. He gave me a call. I faced similar issue when I was working for a different customer. So I was quickly able to identify issue.   Checked if there was any zombie process running (ps aux | grep Z). Yes, there was a zombie process. It was server process. Looks like serve...