We hoped we would not be writing one of these so soon, but we believe above all else in transparency when it comes to a service you choose to trust with your data, and would expect no less when it comes to a customer-impacting incident in our day job, and we would rather make this announcement now on our own terms than have it be discovered publicly 3 months from now.

jump to tl;dr

We developed Sheaf initially as a weekend project exploring the possibility of a Simply Plural replacement. With this in mind, before our specific model of exactly what we were going to build was finalised, we built hooks for several scaffolded features in the backend that have not yet been implemented (for a non-incident-relevant example, member privacy settings - as we currently do not provide any public profile mechanism, these exist but do nothing), and one of these features was the ability to set retention limits for front history. This was based both on a "what tools might someone running their own semi-public instance want available?" and on Simply Plural's shutdown announcement; specifically the mention that front history was a major driver of costs and infrastructure difficulties for them. When we launched Sheaf, we did so with this feature enabled-by-default. We had intended to give ourselves 90 days to look at real usage patterns and determine what would be sustainable for database performance before we had to make a decision, and with the knowledge we could just extend it if needed.

This evening, we noticed when digging into some of the other backend jobs that we had front entries being removed by the cleanup job when we did not expect them. It turns out we made a mistake, and applied the retention cleanup based on the start date of the entry, instead of the intended creation date of the row itself.

If you had any non-ongoing front history entries with a start date over 90 days ago, they have likely been deleted. Rerunning your import will skip already-existing entries and backfill the missing ones, however, if you have edited the start or end times of any entries it could result in a duplicate with the original times being created. If you no longer have your import data available, please DM us on Discord or email [email protected] and we can work with you to restore your history from our database backups.

In hindsight, we were way too cautious with front history - SP had cited front history as being some of their most difficult to manage data and biggest drivers of costs, and from that we had expected and planned for a much higher creation rate on history entries, but from what we have seen so far, we went too cautious on this, and this was more of a reflection of the state of SP's infrastructure than an actual indication that history size was a problem in objective terms, and real projected usage data has not shown this at all, to the point we feel confident we can absorb much larger histories than we initially believed.

We're sorry this happened. This month has not been an easy one for us even outside launch pressure, with major disruptions in our personal life, our own health issues, and a critically unwell dog, and we do feel like we would normally have caught this. Our users deserve a safe home for their data, and we understand the importance of history and will be prioritising a number of followups detailed below. This work should not have been deferred until after launch, and we accept responsibility for this. As in our professional lives, we believe in openness when it comes to incidents and that this writeup is the least we can do.

Immediate remediation:

  • Retention enforcement has been disabled pending a complete rework.

Long-term, before any retention-style feature returns:

  • A retention test fixture with a wide spread of sample data (long fronts, imported history, open fronts, edited times) so any future change to retention logic is provably limited to its intended scope.
  • Clearer in-app communication about where limits or potential limits apply.
  • Nothing removed silently - notification prior to automated cleanup applying, and any affected records surfaced to the user so it's visible and auditable, not a background job only visible in the admin panel. This work in particular was always planned before allowing any cleanup based on retention, but should not have been deferred in the first place.
  • User-downloadable archives of any data affected by retention allowing archiving and/or reimport
  • A shift in how limits apply to important historical data such as front entries, to acting primarily reactively on outlier, potentially-harmful excessive usage.

We apologise for shipping an unfinished feature, and hope these followups will restore some user trust. Much like every other career, engineers can make mistakes, and in our opinion, it is how they handle those situations that makes a truly good engineer - if you can't talk about your mistakes, you can't learn from them; if you cover them up instead of telling people who were affected, you lose infinitely more credibility if someone discovers your coverup months later. As in many things in life, honesty is the best policy here.

If you have any further questions, all the usual channels are available.

tl;dr

tl;dr: history entries that began over 90 days ago (and have since ended) may have been lost due to an implementation bug in a feature we deferred work on. Reimporting your data will repopulate them; please DM us on Discord or email support if you no longer have your import file available