Skip to main content

Can Computers write Science Books? - Brian Clegg

The German academic publisher Springer has for some time been using automated editing software (with mixed results) - but recently has brought out a whole book written by a piece of AI software called Beta Writer. The book, Lithium-Ion Batteries: a machine generated summary of current research, can be downloaded free of charge as a PDF. But is this a serious challenge for science writers?

It's certainly interesting. If I'm honest, this is hardly a book at all - it's more the output of an automated abstract generator pulled together in book form, where frankly this information would be far better just as a web page. However, there's no doubt that there is some interesting work going on here, particularly in the introduction and conclusion sections of the 'book'.

The whole thing starts with a (human written) preface explaining the technology - by far the most readable part of the text. We then get four 'chapters' of machine-generated content, which each have the format introduction/ set of abstracts / conclusion. Obviously it's the introduction and conclusion that provide the most interest.

I'll focus on the first introduction, though the same criticisms apply throughout. The first test of a piece of scientific writing meant to be readable is to take a step back and get an overview of a chunk of text - does it look like English or is it dominated by acronyms and numbers? A chunk out of the first page shows that this is very dense technical text, extremely low on readability:



The other two significant indicators of readability are whether the text is a collection of fact statements or is written using connectives and summary to give flow, and whether or not overall there is a structure that takes the reader by the hand and leads them through a communication process. On both tests, the book falls down in a big way. Pretty well every sentence is a standalone fact statement that could be a bullet point: there is no flow whatsoever. And although some attempt has been made to group these statements effectively, there is no sense of a thought-through structure. In the interminable-seeming introductions - the first one runs to 22 dense pages - there is no sense that we are going anywhere, just that we are experiencing randomly thrown together bits of data.

Inevitably, an automated process will produce some sentences that don't quite work, so one essential here is to see whether these have been captured and fixed. A reasonably high percentage of the content does make grammatical sense, but there are regular hiccups - for example we get: 

  • 'That sort of research's principal aim...' - it should be 'principle' not 'principal'. 
  • 'Materials, a number of metal oxides with high theoretical capacity have aroused more and more attention including...' - that 'Materials,' start makes no sense.
  • 'Through Tang and others, mesoporous nanosheet is synthesized...' - sounds painful.
  • 'It is still maintained the huge capacity of 611 mAg-1... when utilized as an anode.' - doesn't make any sense.
  • 'Apart from, few-layer nanosheets enhance a fast insertion...' - apart from what?
  • And so on for many, many more examples.

Going on comments I've had from some Springer authors, the level of uncaught or automatic-editing-generated errors is fairly high in their human-authored publications - these books tend not to be heavily edited - but because they are starting with far more readable text, this is less of an issue.

So, should science writers be worried? Obviously, as a professional writer myself I'm biassed, but I would say 'No' - at least, not yet. The text in the introductions and conclusions is nowhere near the readability of a decent technical science book, let alone the far higher writing quality required for a good popular science book. And the outcome also emphasises that even if, long-term, automated writing becomes more common, it is always likely to need a look over by a human editor to avoid errors creeping in. However, this is a fascinating experiment and Springer should be congratulated for getting this far.

Comments

Popular posts from this blog

Mathematics with Love – Mary Stopes-Roe *****

Admittedly it’s early days (this review is written in January), but this, for me, is the surprise hit of the year so far! I approached this book with trepidation, but found it absolutely delightful. It is described on the cover as the “courtship correspondence of Barnes Wallis, inventor of the bouncing bomb”, and contains a series of letters between Wallis and his cousin and eventual wife Molly Bloxham, along with some useful annotation by their daughter, Mary. The courtship itself is not without difficulties, as Wallis was 18 years older than the 17-year-old Molly at the start of the correspondence, and her father, not surprisingly, wasn’t too pleased about the interest of such an elderly suitor, but that isn’t the only reason the letters are interesting – it’s also because of maths, and Wallis’s position in the UK as the engineering hero of the Second World War. (Incidentally, it seemed very strange to see letters addressed to “Barnes” – I had always assumed Barnes Wallis was a ...

Data Empire - Roopika Risam ****

The central thesis presented by Roopika Risam is that information gives us (and particularly countries) the power to organise, control and dominate others. Although I have a couple of issues with the presentation, this is a genuinely interesting trip through the history of our use of stored information from the earliest tallies to the latest information technology. I loved a quote from Lisa Gitelman that data is is always 'cooked' so 'raw data is an oxymoron'. This neatly underlines Risam's thesis that data and information are not neutral facts, but rather tools that (like everything from fire to electronics) can be used for good or evil. As we are taken through the historical context, it can sometimes be a little difficult to judge whether Risam regards a particular example as bad or good, even when the outcome is disastrous. One thing I didn't like too much is the adherence to a popular science writing approach that has got distinctly hackneyed: opening chapte...

Andrew Jaffe - Five way interview

Andrew Jaffe is professor of astrophysics and cosmology at Imperial College, London and director of the Imperial Centre for Inference and Cosmology. His new book is The Random Universe. Why science? I’ve always been interested in science, in particular in  astronomy, astrophysics, and space. One of my earliest memories - back in nursery school in New Jersey, I think - was watching one of the moon launches. I wanted that excitement to be part of my life! I never got to be an astronaut, but I did get to be part of the Planck Satellite team, and was privileged to be able to travel to the ESA Spaceport in French Guiana to watch the launch.  In between, I was lucky enough to have a supportive family, get a good education, and find inspiring teachers, mentors, and collaborators. They helped me model the universe, and helped me learn how to refine those models in the face of experimental and observational evidence. That is, they taught me to be a scientist. Why this book? The Random...