Apache Beam number of times a pane is fired with early triggers

我的未来我决定 提交于 2019-12-13 17:43:11

问题


In a streaming beam pipeline, a trigger is set to be

Window.into(FixedWindows.of(Duration.standardHours(1)))
              .triggering(AfterWatermark
                            .pastEndOfWindow()
                            .withEarlyFirings(AfterProcessingTime
                                    .pastFirstElementInPane()
                                    .plusDelayOf(Duration.standardMinutes(15))))
              .withAllowedLateness(Duration.standardHours(1))
              .accumulatingFiredPanes())
  1. If there's no new data between the early firing (15 minutes after the first element of the current window) and the watermark, will there be another firing at the end of the watermark?

  2. If yes, under the same scenario, will there be another firing at the end of the watermark if accumulatingFiredPanes is changed to discardingFiredPanes?


回答1:


  1. Yes. There should always be a firing when the watermark passes the end of the window. The early firing panes will be marked as early, and the watermark pane will be marked as on time.

  2. Yes, currently we always guarantee an on_time pane, which means there will be a firing at the end of the watermark.




回答2:


For #2, you can set the Window.ClosingBehavior as second parameter to withAllowedLateness. There are two variants:

  • FIRE_ALWAYS
  • FIRE_IF_NON_EMPTY

See https://beam.apache.org/releases/javadoc/2.6.0/org/apache/beam/sdk/transforms/windowing/Window.ClosingBehavior.html



来源:https://stackoverflow.com/questions/47760273/apache-beam-number-of-times-a-pane-is-fired-with-early-triggers

易学教程内所有资源均来自网络或用户发布的内容,如有违反法律规定的内容欢迎反馈
该文章没有解决你所遇到的问题?点击提问,说说你的问题,让更多的人一起探讨吧!