Compare the results, e.g. HC5 and quality benchmarks, derived by the new method and the traditional method, and discuss the environmental reality of the results

7 Responses

  • RESOLUTION

    Used Rick’s set of words with modification to incorporate Graeme’s suggestion. Changed Rick’s:

    “Although outside the scope of the current review, a systematic comparison of the various SSD tools would be a useful exercise.”

    To:

    Although outside the scope of the current review, we hope to present the results of a comprehensive review of features together with detailed performance comparisons in a follow-up paper.

  • Agreed that this is out of scope. We could also reference Schwarz and Tillmanns (2019)

  • You could write a paper on this. I think at this stage we simply indicate that comparisons have been made and clear improvements to the data fit in the SSDs have been achieved in many cases using ssdtools Shinyapp over say Burrlioz, and the derived HC5 values differ, whereas in examples where the fits by both are good, HC5s are similar. This will be part of a later evaluation paper.

    • I agree that a systematic comparison is beyond the scope of this ms. I have suggested the following text in the para that spans pp. 8-9, as follows (in bold italics):

      “While we are not advocating adoption of a single standard approach or tool, we think there is a need for closer jurisdictional collaboration, greater harmonisation of methods, and development of at least some benchmark data sets and reference results. The last of these is particularly pressing given the frequency with which we have observed noticeably different HCx values for the same data set from the different tools in Table 1. Although outside the scope of the current review, a systematic comparison of the various SSD tools would be a useful exercise. Some differences between the outputs of different tools are to be expected if different estimation strategies are employed (for example maximum likelihood versus method of moments or single SSD versus a model-averaged SSD) but all things being equal, all tools should give the same point estimates to within some nominally small tolerance (e.g.1-2%). Certainly, differences of a factor of 2 or more are indicative of flawed coding and/or numerical instabilities and convergence issues. The use of ‘reference data sets’ is not a new idea; they were commonly used in the early days of statistical computing to allow both software developers and end-users to assess the adequacy of numerical routines underpinning routine analyses such as ANOVA, regression, and correlation. Even today, the National Institute of Standards and Technology still maintains a number of statistical reference data sets at https://itl.nist.gov/div898/strd/index.html, including the famous Longley data set (Longley 1967).”

      • Thanks Rick. I think we should also mention that we (Rebecca!) will be preparing a follow-up paper that will focus on performance evaluation and case studies.

Comments are closed.