<!-- Benchmark artifact — AROS Skill Library -->
# Benchmark review: PeerJ 20995 (generated by the academic-peer-review skill)

> **What this is.** The complete, unedited referee report produced by the
> [academic-peer-review skill](https://aros.bs-gou.com/skills/academic-peer-review) on
> **2026-08-12**, running **blinded** on the original submission (v0.1) of an article whose
> full peer-review history is public. It is a capability demonstration, not a referee report
> submitted to any journal. The published article incorporates revisions that answer many of
> these points — that is exactly what the benchmark measures.
>
> **Compare it yourself:**
> - Published article: https://peerj.com/articles/20995/
> - Full review history (4 referees + editor decision, CC BY): https://peerj.com/articles/20995/reviews/
> - Scored comparison incl. what the skill missed: https://aros.bs-gou.com/skills/academic-peer-review/benchmark

---

# Review — PeerJ manuscript 2025:08:123854 (first round)

## 1. Basic reporting

The manuscript is written in clear, professional English and follows the standard IMRaD structure; the introduction sets up the problem adequately and the literature coverage is appropriate for a methods-optimisation study. Reagents, antibodies and catalogue numbers are reported in commendable detail. Two located reporting problems need correction: an internal disagreement between the flow-cytometry percentages quoted in the text and those in the Figure 5 legend (concern 5), and a figure-callout error (concern 7). I could not assess figure resolution, panel integrity, or the readability of gene labels from the text extraction alone; these should be checked at print size by the handling editor (**Not assessable from the supplied material**).

## 2. Experimental design

The animal work is ethically approved (approval BDKQ-202504300593, lines 105–106) and the model is appropriate. This is a primary in vitro/animal methods study, for which no formal risk-of-bias instrument is in routine journal use; the relevant reporting standard is ARRIVE 2.0, and its sample-size and blinding items are the ones left open here. Sample size is justified by cell-yield estimates but no power calculation is reported, and randomisation/blinding of the quantification are not described. The substantive design concerns are the outcome-dependent selection of fields for quantification (concern 2), the ambiguous seeding-density metric that is the paper's headline variable (concern 3), incomplete statistical reporting (concern 6), and limited data deposition (concern 8).

## 3. The validity of the findings

The within-method density optimisation is internally consistent and supported by convergent readouts. However, the two claims that give the paper its novelty — that Method 2 yields the *most efficient differentiation*, and that *precursor purity* is important — are not established by the analyses presented (concerns 1 and 4). These are the concerns that decide the manuscript's fate.

## General comments for the author

This study compares three protocols for generating osteoclasts from mouse bone-marrow monocyte/macrophage precursors (BMMs). Method 1 induces osteoclasts directly from freshly isolated BMMs; Method 2 first differentiates BMMs into bone-marrow-derived macrophages (BMDM) with M-CSF for 72 h before RANKL induction; Method 3 adds a Ficoll-Paque density-gradient step before the same two-step induction. Using male C57BL/6 mice, the authors assess differentiation across a range of seeding densities by TRAP staining, resorption-pit assay, F-actin ring staining, RT-qPCR and Western blot for Trap, Ctsk, Mmp9 and Atp6v0d2, and characterise the precursor populations by CD11b/F4/80 flow cytometry. They report that osteoclast formation is maximal at 3–5 × 10⁶ cells/mL for Method 1 and at 1–2 × 10⁵ cells/mL for Methods 2 and 3, with reduced formation at higher densities. Flow cytometry shows the highest CD11b⁺F4/80⁺ precursor proportion and live-cell fraction for Method 2, and no significant gain from the additional Ficoll step in Method 3. The authors conclude that M-CSF pre-differentiation (Method 2) is a simple, effective route and that seeding density is the dominant determinant of differentiation efficiency.

The study addresses a genuinely useful, practical question for laboratories that culture osteoclasts from mouse marrow, and its strengths are the breadth of orthogonal readouts (function, cytoskeleton, transcript and protein) applied consistently across a wide density series, together with clear reporting of ethics and reagents. The within-method density optimisation is convincing. Its main audience is the bone-biology and osteoimmunology community setting up in vitro osteoclastogenesis; the contribution is incremental — the density-dependence of osteoclastogenesis is already documented (Cheng et al. 2022; Remmers et al. 2023, both in the authors' own reference list) — but the consolidated, side-by-side protocol guidance is a fair and worthwhile addition if the comparative claims are placed on firmer footing. As it stands, the two headline claims that distinguish the paper are not supported by a direct comparison and are partly contradicted by the authors' own analysis, so I recommend **major revision**, achievable with reanalysis of the existing data and reframing rather than new experiments, before the manuscript is suitable for publication in *PeerJ*.

### Major Concerns

1. The central comparative claim — that Method 2 provides "the most efficient osteoclast differentiation" (Abstract, lines 40–41) and that the method is a "critical determinant" (Title) — is not supported by any direct statistical comparison between the three methods. Every differentiation readout (TRAP, resorption, F-actin, RT-qPCR, Western blot; Figs 2–4) is analysed only *within* a method across densities, against a control, by one-way ANOVA with Dunnett's post-hoc (per the figure legends). The three methods are never placed side by side on a common differentiation outcome, and Method 1 is tested over a different density range (1–7 × 10⁶ cells/mL) than Methods 2–3 (5 × 10⁴–8 × 10⁵ cells/mL), so "efficiency" is not comparable across them as the data are presented. Could the authors compare the three methods head-to-head on a matched readout — for example TRAP⁺ osteoclast number/area at each method's optimal density — with an explicit between-method test (e.g. a method × density two-way ANOVA where the ranges overlap, or a one-way comparison of the three methods at their respective optima), or otherwise narrow the ranking claim in the title and abstract to what the within-method data actually support?

2. Quantification appears to be drawn from deliberately selected maximal fields: for TRAP staining the "six fields of 10× magnification with the highest number of osteoclasts were selected for imaging" (lines 195–196), and for resorption "the three largest areas with resorption pits were selected for imaging" (lines 204–205). Because the selection is outcome-dependent, it biases the counts upward and weakens any comparison between densities or methods, and no blinding of the assessor is described. Were the fields instead sampled randomly or systematically across the well, was the person performing the quantification blinded to group allocation, and do the reported n (e.g. "n = 6, 10× views" in the Fig. 2 legend) represent independent biological replicates or multiple fields taken from the same well/mouse?

3. Seeding density — the paper's headline variable — is expressed throughout as cells/mL (e.g. lines 169–170, 179–180), yet cells were seeded in "24-well or 12-well plates" (line 178) without stated medium volumes, and the biologically relevant quantity for fusion-dependent differentiation is surface density (cells/cm²). As written, a reader cannot reconstruct how many cells per unit area were plated, nor whether the density optimum is comparable across the two plate formats or across methods. Could the authors report the actual seeding density in cells/cm² (or, equivalently, cells and medium volume per well), and state which plate format and which assay each density series refers to?

4. Two headline conclusions rest on a non-result. First, Method 3 (Ficoll-Paque) is judged "redundant" because CD11b⁺F4/80⁺ purity did not differ significantly from Method 2 (87.4% vs 85.6%, n = 3; lines 431–437) — this treats a non-significant difference from a small sample as evidence of equivalence, which it is not. Second, the Conclusion states the study "highlights the importance of precursor cell purity and seeding density" (lines 46–47), whereas the Discussion's own finding is that differentiation "was most strongly influenced by plating density rather than initial precursor purity" (lines 469–470) — the two statements point in opposite directions. Could the authors soften the Ficoll claim to "no detectable difference" (optionally supported by an equivalence-testing or power argument), and reconcile the Conclusion's "purity" language with the density-driven result the body reports?

### Minor Concerns

5. The flow-cytometry percentages disagree between the Results/Discussion text and the Figure 5 legend: Method 1 is 70.6% (line 417) in the text but 71.3% in the legend; Method 2 is 87.4% (line 423) vs 86.9%; Method 3 is 85.6% (line 433) vs 83.8%. Please reconcile these paired values so each population percentage is reported consistently between text and figure.

6. The statistical reporting is incomplete. The Methods (lines 282–288) name only the Brown-Forsythe test and one-way ANOVA: the normality test used to decide "normally distributed data" (line 284) is not named, the non-parametric test applied when normality failed is not named, and the post-hoc tests (Dunnett's, Tukey's) appear only in the figure legends. Please name the normality and non-parametric tests in the Methods, move the post-hoc specification there, and report the F-statistics with exact p-values so the analyses can be fully evaluated.

7. Figure-callout error: line 308 cites "(Fig. 1C)" for the decline in resorption at 7 × 10⁶ cells/mL, but Fig. 1 is the workflow schematic — this should be Fig. 2C. Please also standardise "Fig." vs "Figure" (e.g. lines 316, 359) throughout the manuscript.

8. Only the flow-cytometry data are deposited (Figshare, line 278). Would the authors also deposit the source data for the TRAP, resorption, RT-qPCR and uncropped Western-blot results, and — since only male mice were used (line 106) — note in the limitations that the optimal densities may not transfer to female-derived cultures, given the known sex-dependence of bone-cell metabolism?

### References

Cheng Y, Liu H, Li J, Ma Y, Song C, Wang Y, Li P, Chen Y, and Zhang Z. 2022. Evaluation of culture conditions for osteoclastogenesis in RAW264.7 cells. PLoS One 17:e0277871. 10.1371/journal.pone.0277871

Remmers SJA, van der Heijden FC, Ito K, and Hofmann S. 2023. The effects of seeding density and osteoclastic supplement concentration on osteoclastic differentiation and resorption. Bone Rep 18:101651. 10.1016/j.bonr.2022.101651

## Confidential note to the editor

Dear Editor,

Thank you for inviting me to review this manuscript. I read it in full and provided constructive, actionable comments for the authors, focusing on:

- Whether the comparative ranking of the three induction methods is statistically supported (it is currently inferred from within-method density series analysed against a control, with no direct between-method test).
- Unbiased quantification and the reproducibility of the paper's central density metric (fields were selected for maximal osteoclast content; density is reported as cells/mL without volumes).
- Interpretation of the null Ficoll result and the internal consistency of the stated conclusions (a non-significant n = 3 difference is read as equivalence, and the Conclusion's emphasis on "purity" contradicts the Discussion's density-driven finding).

My recommendation is **major revision**. I want to be clear that this is not a rejection: every concern is answerable by reanalysis of data the authors already hold (a between-method comparison at matched densities; re-quantification from randomly or systematically sampled, blinded fields) and by reframing the two overstated claims — no new cohort or new experiment is required. If the authors decline the between-method comparison, the honest alternative is simply to narrow the title and abstract to a within-method density-optimisation study, which the data fully support.

My comments concern the study design, statistics and internal consistency; I took the flow-cytometry gating and immunophenotyping at face value from the text and did not independently assess it. I saw no sign of misconduct — the text-vs-legend numeric mismatch reads as a transcription error rather than anything more, and requesting the source tables and uncropped blots (concern 8) is the appropriate, neutral way to close it.

Best regards,
Reviewer
