Jump to content

What would science look like if it were invented today

From IdeaWazaWiki
Revision as of 04:29, 15 June 2009 by wikademia>Mietchen (brought over from http://etherpad.com/XPPme49avV)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)

Feel free to join in to improve this blog post-to-be


Sciengineering: What would science look like if it were invented today?

or: Implementing the wave protocol - possible impacts on science

(to be accompanied by a time slider demo, similar to http://etherpad.com/ep/pad/slider/13sentences but highlighting the collaborative aspect)

Sure, it is hard to imagine you reading this blog post in a world which hadn't yet engaged in science but in the wake of the recent presentation of the Wave protocol addressing the question "What would email look like if it were invented today", it seems reasonable to entertain some similar ideas on reinventing science today, i.e. on designing a system that creates and structures knowledge in a way that both processes can effectively feed on and adapt to each other, making use of the most appropriate technologies at hand.

How can this challenge be tackled? Certainly, it would be helpful to have operational definitions of the terms involved, particularly science, knowledge and research. However, such endeavours are the subject of ongoing debate, so it may be better to skip this step and to design our system such that it would be operational across a broad range of definitions and conceptualizations (for instance, independent of whether or not the term "science" encompasses maths, engineering, medicine, social sciences or the humanities).

Let us start by considering knowledge creation -- or research, for short. Within the framework of existing knowledge, this requires, as a first step, the identification (and perhaps further characterization) of a gap to be bridged or closed (some methodologists prefer or even have to construct their bridges before choosing a suitable place to install them, but we shall not discuss these special cases here).

Once such a gap has been identified (we will leave a detailed consideration of this process to a later post in this series), three basic components (roughly following each other as stages of a research project) are necessary to close it: (1) Planning: an idea on how to bridge or close the gap (2) Realization: the means to put the idea into practice (3) Verification: independent assessment of the realization.

In this enumeration, we have consciously left out a fourth component, very prominent in contemporary science: the communication of selected details of the approach to other members of the scientific community (usually separately for each component - as grant proposals, publications, and control experiments in related studies, respectively). The decoupling of this fourth component from the other three, however, is simply a trait inherited from the era of paper-based scientific communication, and not a technical necessity today when such information can be shared instantly (with few exceptions concerning sensitive information, e.g. patient data) within and beyond the scientific community. For our purposes, we will thus reframe the concept of putting ideas or results on paper as putting them on a wiki, a blog, a dedicated online repository or successors of these (e.g. as blips or wavelets within the proposed Wave protocol).

In this kind of framework (henceforth Public Research Environment, or PRE), individual contributions can be automatically assigned a unique identifier (revision number in wikis, henceforth contribution ID) and linked to its originator (usually the user name, henceforth contributor ID). Currently, the contributor ID is generally unique within but not across individual online platforms (this differs from the paper-based system in which contributor ID is mainly based on an author's surname plus some representation -- variable across journals -- of given names, such that a single contributor ID may be shared by different individuals whose names are identical or similar, while some individuals -- especially those with multiple initials, with non-English characters, or who changed their name after marriage -- may have more than one contributor ID). For online platforms, a number of solutions towards unique identification of contributors have been implemented (e.g. OpenID), including some specifically targeted at scientists (e.g. Researcher ID).

Each contribution ID can not only be linked to its contributor but also tagged (similar to the keywords currently accompanying manuscripts or grant proposals) and have their quality assessed (or rated, for short) by individual contributors (perhaps as a function of the overlap between the tags for their personal expertise and those of the contribution under consideration) according to a pre-defined set of evaluation criteria (e.g. appropriateness to the current stage of a given project, reliability of the information supplied, or presentation with enough context to be understood by specialists and/ or the public). Some journals already allow such ratings (PLoS ONE, for instance) but none of them currently provides aggregations of the ratings by contributor, or incentives to rate (or tag, for that matter) items on their site. In spite of this, the feasibility of generating and aggregating such user-defined metrics has been demonstrated on multiple online platforms, especially in non-scholarly environments (tagging: flickr; rating: Amazon) but also in some scholarly ones (tagging at Connotea). No working implementation exists, however, that would address the lack of incentives for scientists to engage in collaborative research assessment of this sort, but given that funding agencies have managed to coerce scientists and their institutions into all sorts of behaviour during research assessment exercises in the past and present, they should have no problems providing incentives to participate in this one which has the added benefits of being both transparent and beneficial to the scientific community as a whole (it is of note in this respect that there are very few incentives in the current system to deliver timely, fair and detailed peer reviews for grant proposals or manuscripts). One way to do this would be to include both the quality and the quantity of a specific researcher's ratings (both active and passive) into the determination of the variable portion of her re search funding, perhaps with some sort of normalization by the usage frequency of the tags involved (to balance between large and small fields of inquiry, and to avoid exaggerated claims).

The remaining obstacles to a wider adoption of such transparent reputation schemes based on a Public Research Environment with unique contribution and contributor ID schemes are thus not of a technical nature, and we shall assume these features to be available for the system we are about to design. With this in mind, let us now reconsider the three stages listed above:

(1) Planning The conception of ideas is a process very specific to the problem at hand and to the individuals (or possibly even machines) dealing with it. Though technical assistance may be available for a subset of these cases, we will not discuss them here (but in a subsequent post) and assume instead, for simplicity, that a research project is started by being entered into the Public Research Environment and tagged as an idea with suitable keywords. Similar to current systems of journal publication alerts, scientists (and possibly other interested parties, including dedicated robots with their own contributor ID) subscribed to specific tags or contributors (or combinations thereof) will then automatically be alerted of the existence of this new project and may add to it (e.g. comments, references, extensions, limitations, illustrations, links to suitable tools or relevant legal information or related ongoing projects or previous refutations of similar ideas, offers for collaboration or funding, suggestions for a timeline, or simply further tags, or ratings of any of these), to which the original contributor and anyone else interested may respond (with some provisions to avoid spam). No technical difficulties here, just cultural ones associated with the cherished habit of keeping ideas and results private until formal publication.

(2) Realization As a result of these interactions, the planning of a subset of proposed projects will have taken shape, i.e. the necessary material, financial and human resources integrated with a tentative timeline to acquire some preliminary data. Once these are available, they will be posted in the same way as everything before -- with the Public Research Environment acting as an electronic lab notebook -- and immediately integrated with the relevant information available in the system by then, such that the procedures can be adapted as needed to gather the amount and quality of data necessary to bridge the targeted knowledge gap in its most recent state.


more on funding, incl. prizes sorted lists by tags, contributors, ratings, optional: specific tags for specific funders

cloud computing

baseline funding would allow even low-rated ideas to be pursued until preliminary data are available, or further if several researchers combine their respective funds.

(3) Verification It is important to note that such a Public Research Environment would allow for independent verification right from the start in that independent samples could be investigated in parallel by independent scientists (or even robots) following the same public protocol and posting their data in public as they arise - a situation far from being common in contemporary science, although not entirely new after successful completion of large-scale collaborative initiatives like the Human Genome Project.


One of the most frequently raised arguments at this point concerns the perceived danger of being scooped of the information laid out under the eyes of everyone and their dog (or parrot). But with the system described here, it will always be possible to point out, in public, who had posted what and when, thereby severely limiting the effectiveness of any scooping attempt. There is another, probably far more relevant possibility to take into account here: If your idea is public from the start on (and scientists have become used to this way of communication), it is possible that others will provide constructive feedback early on, including comments on your design, offers to collaborate on either of the three basic stages, or suggestions on extending the approach to other knowledge gaps.

Indeed, once adoption of a "posting-in-public" attitude reaches a threshold, the incentives will be on the sides of those who post their stuff immediately (though provisions should be made for some special cases, e.g. concerning the privacy of personal data from human subjects and patients).


OK, you might say, but what about peer review then? --cite examples from Cameron's letter, or better, on it performing only slightly better than chance alone, and given the costs involved, it is certainly not an effective way of quality assessment.

--needs incentives, not present in the current system

--possibility to invest in selected projects initiated by others is perhaps even better a form of assessment than classical behind-the-doors peer review


pop culture effects

--not designed to detect fraud. Nor is wiki-based communication, but it facilitates to detect fraud.


Thus, the little change in design -- switching from paper-based to web-based communication -- may have profound consequences on the three basic components: the permanent communication of progress during the course of the project will shorten the feedback loops, allowing to improve it on the run and linking it to other gap-closing or even maintenance work on our shared corpus of knowledge.


When the whole scientific cycle is open, this has important implications for the media: Instead of creating a stream of "scientists found out" broadcasts, they can add in some of the "scientists are currently investigating - let's see how they do it" variety.

Next issue: What would knowledge structuring look like if it were invented today? including (1) identification of a gap in existing knowledge mention teaching/ outreach


tags: Google, Google Wave, Science, Science history, Future of Science, Peer review, research funding, young scientists, Friendfeed, wiki, scholarly wikis, translation, impact metrics, impact factor, scientific collaboration, internet protocol, open science, grants, science funding, Fantasy Science Funding

Old text bits

More on Google Wave


  • use aspects of the GW protocol for a new and open standard of impact metrics, incl. grant proposals, research funding etc.


The basic steps:

  • Give each research contribution a unique ID and non-unique tags
  • Give each contributor to research a unique ID and the option to tag and rate contribution IDs (perhaps as a function of overlap between the contribution's tags and those on their own expertise)
  • Aggregate the ratings over contributions, contributors, tags and combinations thereof

Open scientists make their work transparent in the world wide web, e.g. by: discussing own research ideas, methods and results with others in the internet keeping a public lab notebook making their teaching open to public discussion blogging about their scientific activity and while doing this not only reflect about their own action, but also encourage others making problems public they are actually working on and thus giving other the chance to participate actively in the problem solving process submitting articles to publication institutions which have set up a public review process assessing articles for publication institutions which have set up a public review process making raw data and analytical tools (on which their works are based) publicly accessible

Q: is there a word for the process of structuring knowledge? Learning? Epistemology?

  • feature request: offline archiving, selective download (ideally in an automated fashion, like with "saved searches"), rating of individual contributions, and analysis of contributions via an adaptation of PageRank

Google Wave

   Federation: decentralized interactions via the wave protocol
   

Google wave and implications for science : http://wwmm.ch.cam.ac.uk/blogs/murrayrust/?p=2032

Google Search Google Scholar Google Earth Google Docs Google Mail Google Chat Google Images Google Knol Friendfeed


Part I: Inventing science funding basis: http://ways.org/en/blogs/2009/may/24/implementing_fantasy_science_funding focus on transparent reputaion system beyond the Journal Impact Factor

Part II: Inventing scientific collaboration basis: http://ways.org/en/blogs/2008/dec/28/the_journal_scope_in_focus_putting_scholarly_communication_in_context (rethink in terms of waves) and the blog3 proposal


Sidenotes: translations, if harvested from a web corpus, will only work well if there is a large corpus containing the relevant phrases, which is certainly a problem for scientific topics




I'm sure you will have noticed, but I feel I should mention it nonetheless: this blog post was inspired by Google who posed the question "What would email look like if it were invented today?" last Thursday and answered it during the same impressive presentation with Google Wave (briefly announced in my previous post) - a tool .

the above paragraph could also go into the lead (rephrased)


  • Google Wave Widgets

http://zope.cetis.ac.uk/members/scott/blogview?entry=20090601115357

  • Six reasons why it may become important

http://thinkvitamin.com/dev/six-ways-that-google-wave-is-going-to-change-your-business-career-and-life/

  • Cameron's initial comments

http://blog.openwetware.org/scienceintheopen/2009/05/30/omg-this-changes-everything-or-yet-another-wave-of-adulation/

http://bjoern.brembs.net/news.php?item.521.3

  • Chicago Sun Times

http://www.suntimes.com/business/1606282,ihnatko-google-wave-060309.article

  • Michael Nielsen: Doing science in the open

http://physicsworld.com/cws/article/indepth/38904

  • Google Wave, by David H Burton

http://howpublishingreallyworks.blogspot.com/2009/06/guest-post-google-wave-by-david-h.html

  • Transclusion gets the wiki close

http://forum.citizendium.org/index.php?topic=2706.msg21441#msg21441

  • Going Beyond a Public Understanding of Science to Give the Public a Voice

http://pathtosustainable.wordpress.com/2009/06/02/talking-science/

  • A series of five blog posts:

http://www.endesha.com/blog/on-google-wave-part-1-architecture/

  • Voting/ reputation scheme implemented:

http://stackoverflow.com/faq

  • 100 Awesome Open Source Tools for Writers, Journalists, and Bloggers

http://www.onlinecourses.org/2009/06/09/100-awesome-open-source-tools-for-writers-journalists-and-bloggers/

  • Cameron Neylon: Head in the Clouds: Re-imagining the experimental laboratory record for the web-based networked world

http://cameronneylon.wikidot.com/head-in-the-clouds-automated-experimentation

  • Poster pre-conference review

http://friendfeed.com/the-life-scientists/ce7887f1/hi-all-please-have-look-at-this-almost-completed

  • organizing meetings via IRC

http://wiki.okfn.org/wg/science/1#

  • Programming: 10 principles that will guide the evolution of scripting languages in the future

http://www.infoworld.com/print/38839

  • TRAFFIC in 2020

http://friendfeed.com/the-life-scientists/41fa74f8/vision-of-traffic-international-journal


A possible fourth component would be the communication of selected details of the approach to other members of the scientific community (usually separately for each component - as grant proposals, publications, and control experiments in related studies, respectively). The decoupling of this fourth component from the other three, however, is simply a trait contemporary science has inherited from the paper-based era, and in most cases not a technical necessity today.

Currently, we do this chiefly by means of publishing journal articles or monographs and by giving talks or presenting posters at conferences.

(1) identification of a gap in existing knowledge

(let us call them "research" and "learning" here)

system whose only input are verifiable observations and whose output is

in which existing and newly created knowledge

in which the creation and maintenance of knowledge are intimately integrated with one another and based on verifiable observation (be these empirical, experimental or theoretical).

any coherently organized system of knowledge attained by verifiable observation or logical vindicatio

that creates and provides knowledge in a way that both processes can effectively feed on each other. integrated

our leaders had judged our current knowledge system to be approximately coherently structured but in dire need of a wholesale revision of its interface to observations: research. They added three more points:

  • The observations should initially feed on the existing knowledge system but easily adapt as the system changes.
  • The knowledge system, in turn, should feed exclusively on verifiable observations, but easily, quickly and transparently adapt as the system changes.
  • You are in charge.


procedural knowledge

The point here is to make use, right from the start, of the most appropriate technologies available on our planet today (ignore, for the time being, where these may have come from, or imagine, for simplicity, that they were given to us by an alien civilization to which we have lost contact).


Carl Linnaeus (1707-1778)[1] [2] established for the first time a widely acceptable and fruitful set of principles for classifying plants and animals into the groupings we know

that the interface should make use of the most appropriate technologies available today, that it should only allow verifiable observations

you were charged with a wholesale redesign (and thus basically the invention) of an evolving structured knowledge system that feeds on and into verifiable observations. Knowledge in this context includes procedural knowledge like drawing, singing or programming skills, and observations may well be abstract, e.g. concerning the behaviour of a mathematical function when it approaches a singularity. Suppose further that an acceptably coherent knowledge system were already existing (ignore, for the time being, where it may have come from, or imagine, for simplicity, that it was given to us by an alien civilization to which we have lost contact) but it is rather static - what is missing is the interface to observation, something that feeds both on and into the knowledge system: research.

Citizendium: Science is "any coherently organized system of knowledge attained by verifiable observation or logical vindication".

Me: the cyclic process of verifiable observation that feeds on and into a coherently structured system of knowledge.


To me, "observations" includes abstract notions like observing the behaviour of a mathematical function when it approaches a singularity.

Definitions: scientists: One who works on scientific principles. http://dictionary.oed.com/cgi/entry/50215801?query_type=word&queryword=science&first=1&max_to_show=10&single=1&sort_type=alpha

science: 5 b. In modern use, often treated as synonymous with ‘Natural and Physical Science’, and thus restricted to those branches of study that relate to the phenomena of the material universe and their laws, sometimes with implied exclusion of pure mathematics. This is now the dominant sense in ordinary use. http://dictionary.oed.com/cgi/entry/50215796?query_type=word&queryword=science&first=1&max_to_show=10&single=1&sort_type=alpha