Ops incorporation - #212
Wesley-J-Davis wants to merge 5 commits into
Conversation
…y launches the code and maintains logs
…res actually fail the jobs in D_BOSS, since the perl driver is listening for a bad return from the python.
|
Just to confirm the intended behavior here: I agree that we want actual preprocessing failures to propagate a non-zero exit status so the ops job does not appear successful. There are two cases where this changes the previous behavior from continuing to aborting the entire job:
Current Also, the current log message says
Current If both cases are supposed to make the ops job fail, then this change looks good to me as well. |
The original code intended to skip the date and continue processing. I'm not sure what the preferred behavior is in the OPS setting.
I think the preferable setting here would be to skip the bad file and continue. |
I will work on a solution to that effect. This is a good catch, we in ops also need to reserve the ability to force processing even if there'a bad granule. I think we would prefer the default behavior in case of bad processing of a granule to fail the entire job, but yes we do need to be able to force it through. |
|
I am having second thoughts on these commits. I think I would be cleaner to have an operational version called We would then keep the default version of Otherwise it will require amending the What I would be doing is just taking the proposed changes to SMOSproc_main.py and applying them to SMOSproc_main_ops.py and switching between the two in the perl driver that we will be using to connect D_BOSS to the submission of this job, which is not a part of this repository. How does everyone feel about this? |
|
Others may have other ideas, but here is my 2c. My only concern would be duplicating the full Could we instead keep That would also avoid having to modify |
This is a good idea, thank you. Here's my plan to incorporate it, slightly different than suggested: In the perl driver I set an environment variable that reflects whether we want to force the job through or not: In SMOSproc_main.py I read the environment variable to determine which behavior path to choose upon failure: Then later on down in main() I put the if-else blocks in to either hard exit or continue. This makes changes to the repository minimal, doesn't add new files, and allows ops to change the behavior of the code dependent upon our DBOSS configs. How does everyone feel about this method? |
|
Yes, I like this approach. It keeps a single implementation, avoids adding another entry point, and lets the Perl/D_BOSS driver select the error behavior for each run without changing get_time_range(). The only thing I might change is the environment variable parsing so an unexpected value doesn't automatically become True. For example, we could explicitly accept true/false and either default to False when it is unset or raise an error for any other value. Otherwise this looks good to me. |
…w processing steps as they happen with more fidelity.
Wesley-J-Davis
left a comment
There was a problem hiding this comment.
I used a try/except block to check for the value of the FAIL_FAST environment variable. If it is set incorrectly, or unset, it defaults to the original behavior of the code.
Testing has revealed that these settings do indeed control the flow of the code as intended.
I added one logging command to run_in_isolated_process so that it's initiation shows up in the logs, which helped me follow what was happening.
|
It looks like the main while current_date < end_time: block was accidentally duplicated after the first loop.
|
|
Do we have a way to be alerted if there is no input data for a given date? We normally expect 20+ zip files per day. |
Yes. In |
We pretty much always expect 20+ files. If fewer than that, we need to take a look. |
The DBOSS configuration of this job requires a successful download of the SMOS data that is being preprocessed. That job is known as GET-SMOS-01. So, without a successful download of all three types of SMOS data that I'm downloading (BWLF1C,SCLF1C,SMUDP2), this job won't run at all. This is at least a preliminary check, but we don't always get 20+. Yesterday we got 28 for this product and today we got 2. I'll plug in a check before the ee to nc conversion step to look for a minimum number of files, sounds like we've settled on 20 as our target? |
It looks like more files are still arriving based on the timestamps, so let's check back this afternoon to confirm today's total. Our estimate of 20 files is conservative based on past counts, but we'll adjust accordingly if upstream processing changes. |
… beingallowed to continue
|
Inserted the following change to check for minimum of 20 ee files before conversion to nc. |
Small modifications to SMOSproc_main.py, which cause any failure in the preprocessing, or reg2fit steps to return a non-zero exit code that can be detected by the perl driver that we in ops use to run these scripts and return their outputs to listing files.