Oracle MFT Error - Error occurred while polling the listening source
We can setup email notifications whenever an error occurs in MFT like failed connection to source or target etc. This is very useful feature of MFT. We had setup this monitoring for our MFT servers. We used to get below error frequently in production server while connecting to a 3rd party application.
Error occurred while polling the listening source XYZ
Cause: Unexpected error
occurred while polling the JCA source.
Action: Review the diagnostic log for the exact error, which begins with the
string 'ERROR:'. This error line, along with the preceding lines, indicates the
reason for this exception. Contact Oracle Support Services if the problem
persists.
The endpoint here was the target system to which MFT transfers file. We had setup retry at target so transfers were retried whenever this error occurs. So, it didn’t have direct impact on business. But we used to get around 100 alerts mails with these connection errors. It was tiring experience to read and check if this was real alert or false alarm.
We had setup a meeting with 3rd party application vendor and explained our scenario. They asked for sometime and came back that they didn’t find any issue from their end. Issue stopped after some time and all were happy. Unfortunately, it reared its head again after couple of weeks. We approached vendor and they told same thing. Issue used to exist for weeks and go away for weeks together.
We didn’t have direct contact with 3rd party vendor. It had to go through SME. Our SME for this area was known for perfection and he didn’t want to give up. He escalated to higher levels and everyone got into a call. We ran sftp commands on the call and around 30% of the sftp connections were failing. Which was same as failure rate we see normally. So we concluded that there was no issue with MFT.
With the information in hand and with experts on board, the vendor did real analysis for the first time. They came back and told us that they saw lot of failed login attempts to their test env. There was some common gateway hardware between test and production systems. These failed connections to test envs impacting common gateway and hence impacting production system too.
We were surprised that such a reputed vendor had such a setup. The troubling transfers were setup during POC stage and passwords got changed later. We should have cleaned these transfers immediately after POC. Anyhow, it was good lesson to everyone. We didn’t see these errors again after correcting passwords.
One question is still not answered, why issue used to absent for days together and reoccur again? The reason was test MFT servers used to go down due to OOM error (bug in MFT) and we used to restart them only when explicitly requested.???
Comments
Post a Comment