Skip invalid timeframes during ROOT input - #15576
Conversation
|
thank you for your contribution. somehow i did not get notified about this. i will have a look at it tomorrow, time permitting. One comment i already have is that you can probably split the changes to LifetimeHolder as a separate PR. |
da4d742 to
404624e
Compare
bb69cda to
d369614
Compare
2f0f392 to
39c7515
Compare
|
The part about the MessageContext can also be merged separately. Can you spawn a new PR? |
|
@ktf Do you have other comments apart from the splitting? |
|
I did not look into the details of the rest. I will do once I am back, tomorrow. |
|
We should also think how to test this, e.g. making sure it actually does not get stuck or silently skips stuff. |
Once we can enable it by environment variable, we plan to clone some trains which had large number of failures, enable the feature only for those, and do a detailed inspection of the logs and output. |
| if (!treeRead) { | ||
| if (!first) { | ||
| LOGP(fatal, "Can not retrieve tree for table {}: fileCounter {}, timeFrame {}", concrete.origin.as<std::string>(), fcnt, ntf); | ||
| throw std::runtime_error("Processing is stopped!"); |
There was a problem hiding this comment.
Does this actually trigger? I suspect this is old dead code.
There was a problem hiding this comment.
My understanding is it can still trigger if a later requested table uses a separate DataInputDescriptor whose file has fewer timeframes than the first table’s file, but for the normal case this is effectively unreachable. Should this be removed?
Handle corrupt reads as recoverable and discard the affected timeframe when DPL_AOD_READER_SKIP_INVALID is enabled.
cac116c to
a9a11c7
Compare
| }; | ||
| auto readState = TFReaderState::READ_FIRST_TABLE; | ||
| size_t routeIndex = 0; | ||
| while (readState == TFReaderState::READ_FIRST_TABLE || |
There was a problem hiding this comment.
Sorry, but I find this still hard to read / understand. I would have expected the state machine to work as follows:
while(!exitStates(state))
{
case STATE: {
// Perform operations for state
state = NEXT_STATE;
} break;
case STATE2: {
} break;
// etc.
}while here it's all a bit intermixed. Am I missing any particular corner case which makes this complicated, hence the current approach?
There was a problem hiding this comment.
I've pushed a change now to try clarify and separate this
| MetricSpec{.name = "dropped_computations", .metricId = static_cast<short>(ProcessingStatsId::DROPPED_COMPUTATIONS), .kind = Kind::UInt64, .minPublishInterval = quickUpdateInterval}, | ||
| MetricSpec{.name = "dropped_incoming_messages", .metricId = static_cast<short>(ProcessingStatsId::DROPPED_INCOMING_MESSAGES), .kind = Kind::UInt64, .minPublishInterval = quickUpdateInterval}, | ||
| MetricSpec{.name = "relayed_messages", .metricId = static_cast<short>(ProcessingStatsId::RELAYED_MESSAGES), .kind = Kind::UInt64, .minPublishInterval = quickUpdateInterval}, | ||
| MetricSpec{.name = "aod-invalid-read-skipped-timeframes", |
There was a problem hiding this comment.
The extra metric can go in even without the rest of the code, if you want.
| int level = originLevelMapping.empty() ? -1 : 0; | ||
| auto fileCounter = std::make_shared<int>(0); | ||
| auto numTF = std::make_shared<int>(-1); | ||
| bool const skipInvalidReads = [] { |
There was a problem hiding this comment.
This can probably be a static function, no?
There was a problem hiding this comment.
Also, didn't you have some more complete logic for it? what happened to it?
There was a problem hiding this comment.
There wasn't a more complete logic than this, I've also checked force pushed I made previously to make sure. What way should this better be handled?
| try { | ||
| f2b->fill(datasetSchema, format); | ||
| } catch (std::exception const& e) { | ||
| f2b.discard(); |
There was a problem hiding this comment.
Maybe all the exception handling you are doing should be rewritten using std::experimental::scoped_exit (which our compiler requirements support) and std::throw_nested_exception .
|
@ktf We just discussed and think it's best we do the parent file exception implementation later. In a first step we will test this feature only on demand on certain datasets on Hyperloop and we can start testing with non-parent file datasets. |
Handle corrupt reads as recoverable, discarding the timeframe