Showing posts with label databases. Show all posts
Showing posts with label databases. Show all posts

21 November 2025

Google Scholar finally falls to “AI in everything”

Who thinks Google searches have gotten better recently? Because I have not seen anyone say that.

A few days ago, Google Scholar started its version of putting “AI” (large language models) to Google Scholar. Because that’s what every app does now whether users want it or not.

Google Scholar was one of the few online services I used regularly that hasn’t shown signs of enshittification. Scholar just worked. The complaints about it were things like that the metrics could be gamed, and it wasn’t perfect at screening out non-academic content. But I never heard from the online research community that the core search function was somehow deeply deficient at finding relevant papers.

Disappointed. But I expect this from Google now, just like I do from every tech company.

26 September 2017

Microsoft Academic: Second impressions


By happenstance, I thought of Microsoft’s imitation of Google Scholar, Microsoft Academic Research, yesterday. I reviewed it years ago, but hadn’t thought of it much since then. I wondered what happened to it.

Quite a lot, as it turned out. The original website was decommissioned, and the name was shorted to Microsoft Academic. Version 2 launched in July 2017. I thought it was worth a new look.

Sadly, the second look is not much more promising than the first one.

One of the changes is that you can create a profile of your papers. That could be good. I’ve found profiles in other similar sites to be kind of useful. Okay, you have to create an account. No problem, I do that all the time... and you hit the first oddity.

Weirdly, you must use another site to set up your account. You can’t just give an email address and pick a password, like pretty much every other website on the planet. You have to use Twitter, Google, Facebook, Microsoft, or LinkedIn.

I thought, “My institutional ID is handled by Microsoft, and I use that to log in to Office 365, so that should work.” Nope. So I logged in with Google.

I discovered a possible reason why Microsoft Academic won’t let me use my institutional email as I start building a profile. It asks for institution, and snootily insists, “An affiliation must be selected from the list of suggestions.” Except that, according to it, the University of Texas System consists of just five institutions, not fourteen. And my university was not one of those five.


Building out a profile was weird, too. I thought that this next gen scholarly database would support ORCID, so you could enter one number and have it gather all your publications. Nope.

Microsoft Academic seems to identify authors by some mystic combination of name, institution, and... something else. For example, it considers the Zen Faulkes who was at the University of Melbourne a different author than the Zen Faulkes at the University of Texas-Pan American.


So you have to go in and “claim” your papers from however many random ways Microsoft Academic has parsed your name. I have a very distinctive name, and my papers were split into something like ten different authors. I cannot imagine how many different ways publications might be split if you have a common name. Or if you changed your names.

For some of my new papers, I found them by searching the database for the title. I added them to my profile by clicking the paper’s hyperlink, clicking “Claim,” then going to another page and clicking “Submit claim” again. It seemed to be a lot of clicking.

The profile lists “top papers,” but the metric Microsoft Academic uses to determine “top” papers is not clear. It’s not citations, because the number of citations in my list of papers are: 11, 28, 17, 9, 16.

Maybe the profile has a few weak spots, but good coverage of the scientific literature might make it valuable. I searched for “Emerita benedicti” (on my mind since the publication of my newest paper last Friday). That gives 15 hits in Microsoft Academic, but over 2,300 in Google Scholar. But if I search for that exact combination of words in Google Scholar, I am still left with over 32 hits, more than double Microsoft Academic’s yield.

Microsoft makes much of Academic’s “semantic search,” so it may be that I will find it more useful as I try other, more complex searches, rather than simple things like looking for a simple species name.

The home page for all this provides a customized dashboard with a calendar of scientific conference, research news, recent publications, and recent citations of your papers (not visible in the screenshot below).


Google Scholar gives you a couple of alerts on its home page, a sparse approach that leave you with no doubt as to what its job is: Google Scholar is a search engine. It’s not quite clear what Microsoft Academic wants to be. The home page of Microsoft Academic feels more like it wants to be a science news feed; more the science section of Google News than Google Scholar.

Perhaps the kiss of death in all this is that practically everything on this site feels like it’s moving through molasses. It’s sloooooooow. I spent a lot of time looking at a screen like this, waiting for it to populate:



Even while writing this blog post, I got a “You do not have permission to view this directory or page” error message when I tried to go to the home page. Google Scholar feels like it’s using telepathy in comparison. (Update, 26 September 2017: A bot attack may have been responsible for this slow performance.)

I will keep trying Microsoft Academic for a while to see if I learn more. But this project is now at least six years old. And darn it, it still feels like a clunky beta version, just like it did back in 2011, not a modern 2.0 version of an academic search engine that it says it is.

Related posts

Microsoft Academic Research: First impressions

External links

Zen Faulkes’s  Microsoft Academic profile

03 August 2015

Into the vault: oxidative stress and ascidian embryogenesis


I’ve talked before about the long waits in getting projects published. But sometimes, despite waiting, projects never make it past the conference poster stage. I’ve also talked about developing a gut instinct for whether something is publishable.

It’s nice that now, there are ways to turn ephemera into an archival, potentially usable and citable, document. For a while, I’ve been meaning to start putting up some of my posters into FigShare, which I’ve been of fan of from early on. I first used it when I published a paper here on my blog. Since then, I’ve used it to archive the raw data for several of my papers as unofficial supplemental information.

The first one to go up is a poster I presented at the third International Tunicate Conference in 2005 at the University of California Santa Barbara.

This one is one of the relatively few projects that we were never able to push out into a paper. I still think it makes for a pretty good poster, though.

Archiving this poster got me thinking. I see clear value in archiving old posters that can document projects that never made it into the scientific literature. But is there value in archiving posters that were the early versions of projects that did make it into the regular scientific literature? I can see old posters have some interest as examples of design (see the Better Posters blog). They might eventually have some historical interest.

But is there any scientific interest in archiving old posters? Posters are generally works in progress, so tend to be incomplete and preliminary. Might they actually confuse matters by including dead end ideas that were abandoned by the authors?

Reference

Stwora A, Scofield VL, Faulkes Z. 2015. Effects of oxidative stress on Ascidia interrupta embryogenesis. figshare. http://dx.doi.org/10.6084/m9.figshare.1499282

02 September 2011

Microsoft Academic Research: First impressions

Microsoft Academic Research is a new service that continues the company’s attempts to catch up to Google’s services – Google Scholar in this case. I hopped over and typed in my name. Not just out of vanity, mind you, but because I know what should be returned.

The search allows you to select certain domains. A lot of my research is spread over a wide field, so I checked as many boxes as I thought might be even tangentially relevant.


Oh dear.


I got five hits: Three in neuroscience and behaviour, one in biology and biochemistry, and one in clinical medicine. (Wha...?) Even for a beta version of the service, I was expecting double digits at least. Maybe I should have checked more boxes.

And... what a second... who’s that guy?


I have never met Michael N. Nitabach. He wasn’t an author on the paper. A little clicking reveals that there’s an “Edit” button, and that I can remove him as an author:


What else can I do here? Ah, there’s a spot to add a PDF link. Since this paper is open access, I can do that. Easily, let’s grab the DOI and head to the page...

What?

Why is the DOI taking me to Figure 7 instead of the main page?

Okay, let's fix the DOI for the article. And add the link to the PDF. Might as well copy and paste the abstract too while I'm here. Why am I doing this work again? Isn’t the point of the database to have this stuff for you?

Sigh.

Also noticed that the “type” of article includes “poster”.


While I am a big poster booster, I don’t know that I want posters, which are typically very gray literature, in an academic research database. But Google Scholar catches blog posts sometimes.

Back to the main page. Hm. What are these “Were you looking for these authors” bar along the top? Yup, that’s what I was afraid of. Each one has pulled a different set of papers, even though all three of the alternates specify the same name. Why is the search for my name returning four hits instead of one? It’s not as though one was “Z Faulkes” - they all have the same name spelled out, although one includes my institution’s address below it.

The third “were you looking for...” – the one with my address underneath – is interesting, though.


This profile has a nice little dashboard, an RSS feed for updates... still missing a buttload of information, though. Let’s try editing author information. Ah, here is where add pictures, my home page, and - aha! - I can merge the three different profile!

But it’s going to take a week to verify the merger. While I appreciate that I can submit corrections, there is going to be too much literature to crowdsource corrections.

My first impression of this service is to ignore it for about a year, or until I hear about a major update. There’s just too many odd and unpredictable things going on here. It’s not trustworthy yet.

Hat tip to Bjorn Brembs on Twitter.

Additional: Bjorn Brembs followed up, saying:

It's less than beta and some/much of what you mentioned they are aware of.

29 December 2010

The cloud is fickle

Following up on my recent concerns about archiving, stability, and the scientific record, I was disturbed to read this:

Yahoo! is shutting down Yahoo! Video next year. March 2011. March 15, 2011, to be exact. They will delete all user-generated content on that day.

Yahoo! Video was the second-most used video hosting site behind YouTube. Number 2! And all of it, all video, is going to be deleted. Thousands and thousands of videos, many of which are likely hosted nowhere else, completely gone.

With my original post, I was trying to argue that it wouldn't take civilization ending to pose archiving problems. All it would take would be slow, creeping link rot.

The closing of a major video service isn’t quite the collapse of Western civilization. But it’s more than link rot.

03 March 2009

I'm not quitting my day job regardless of this estimate



The value of the NeuroDojo blog, according to $timator:

$4,571,815

Check yours?



Um.

No.

There’s no way that’s even close to reality. My blog is not worth $4,571,815. I agree with Beth’s Blog on this one.

If you are attempting to translate benefits into dollar amounts, particularly intangible benefits - make sure you have a credible formula that you can explain to your executive director in under 5 minutes and that is credible.

By comparison, here’s the estimate for the main page of the institution I work for:





My blog is not worth 100 times more than my university’s main page ($38,413). Even I don’t have that big of an ego to think so. My blog is also worth more than Richard Dawkins’ main site ($187,860), the Pharyngula blog ($2,361,839), Science magazine ($933,301), the National Science Foundation ($242,149), and the government of Canada ($191,781).

I suspect the issue is that because this blog is hosted on blogger, it’s somehow incorporating values for all of "blogspot" in its calculations, inflating the estimate horribly. My Marmorkrebs blog is valued at $4,570,988, but the main Marmorkrebs page at $254.

They try to explain this, but in a terribly confusing way:

For a sub-domain you have to take under consideration if you are in a (for example) tumblr or blogspot context or in a situation where the sub-domain is an important piece of the main domain otherwise the final estimator value could be higher.

Pardon...?

I don’t think the site maintainers’ main language is English.

03 July 2008

Long live the scientific method

Chris Anderson provokes with an article titled, "The End of Theory: The Data Deluge Makes the Scientific Method Obsolete."

There's some interesting ideas, but the argument is based on a false premise.
This is a world where massive amounts of data and applied mathematics replace every other tool that might be brought to bear.
It's perhaps understandable that an outsider, a non-scientist would mistakenly believe this premise to be true: that there are massive amounts of data available for all scientific problems.

There are not.

There are only a few fields of science that generate large amounts of high-quality data. I'm thinking maybe some branches of physics (like nuclear physics, maybe astronomy), social sciences (demographic and census data, automatic tracking of web useage), and maybe genetic data for a select few animals (humans, mice, fruit flies, Arabidopsis).

These are the exceptions.

In most cases, scientists have to eke out by hand one experiment at a time. It's not automated, it's not massive, and it doesn't generate huge numbers. To take an example from my field, invertebrate neurobiology, there isn't really good agreement on how to describe neurons in such as way that they can be put into a searchable database (although the NeuronBank project is making an effort to at least think about that problem).

Anderson goes on to say:
There is now a better way. Petabytes allow us to say: "Correlation is enough." We can stop looking for models. We can analyze the data without hypotheses about what it might show. We can throw the numbers into the biggest computing clusters the world has ever seen and let statistical algorithms find patterns where science cannot.
Scientific theories have three traditional virtues. Predict, control, explain. Massive datasets may indeed give us pretty good predictive power -- correlations often do. It may not give us control. And it certainly doesn't explain. We really need causal mechanisms to explain.

For instance, let's take climate change. If it were the case that massive data is all you need, there would seem to be no need for the ongoing debates about climate change. We have massive datasets there. And indeed, the scientific questions are supported by a large consensus. But people don't care that there's a correlation between carbon output and temperature change, they want to know if one is caused by the other. The policy decisions are very different depending on what your thinking of causal mechanisms are. Cause is king.

Make no mistake, automation changes things. But it doesn't change everything.

26 February 2008

A good day for the world

The Encyclopedia of Life has finally moved past the preview stage.

Good luck at trying to access the page. It's been very slow today, the server no doubt reeling under the load of people who have been eagerly awaiting it.

This particular database spring from a TED prize for biologist E.O. Wilson, shown below.



Not to belittle Wilson's enormous contribution for this project, but I do need to say that kind of project has been on the minds of a lot of people for a long time. There are various taxonomic databases out there. The Tree of Life was probably the first major one, and Wikispecies is another. And those projects have been very valuable, but I think it's fair to say they haven't revolutionized the science the way that GenBank did for DNA or that Wikipedia did for general knowledge.

Hopefully, Encyclopedia of Life can be that transformative resource.

It's interesting to compare how different databases look. Let's take spiny sand crabs, Blepharipoda occidentalis, the main species I work with for my doctorate. In Tree of Life, there isn't a listing for the species or even the genus. Just a species name in Wikispecies. Like the Tree of Life, I can't even get close to "my" sand crab species in the Encyclopedia of Life yet, but I think you get a sense of the ambitious nature of these projects.

31 May 2007

Revisions and wow

It has taken more days than it should, because I have been lazy with no excuses, but I finally resubmitted a rejected manuscript to a new journal today. I am happy to have something in the hands of an editor again. Now, I just have one more manuscript to revise and resubmit, and then I can start real data collection again.

From the always intriguing TED talks, this one about Photosynth is amazing. There is a demo. You must try it to believe it. Very, very cool.

There must be ways to start stitching biological information about species together the way that these guys have stitched together photographic information. Hm. I must think on this some more.

01 September 2006

Google knows I'm typing this

This article picks up on something that has been worrying about. Many academic journals have started to use Google Scholar to link to other articles. There are similar, public databases -- like PubMed -- but Google Scholar does a lot that others don't, probably has more coverage of articles. It's worrying to me, though, that a private company has so much control over our ability to find scientific content. And nobody seems to notice. I think I will have more to say about this later, as it links up with some thinking I'm having about dispersing scientific information.

22 November 2004

Research Google

It's fair to say that search engines have revolutionized how people use the internet, and in fact, much of their lives. And the 900 pound gorilla on the search engine block is, of course, Google: the only search engine whose name has entered the language as a verb. One of my colleagues said to me at a meeting, "I solve all my problems with Google now. 'Daughter annoying'? 'Common problem' comes the reply." (The name of said colleague will not be revealed on the off chance that his daughter Googles his name and hits this blog.)

I've just learned from this article in Nature that Google has put up a scientific version of its search engine called Google Scholar. It's still "beta testing," but usually these test versions work fine.

I bookmarked this page as fast as I could. This is going to be an amazingly powerful work tool. There are other science related search engines, chief among them Pub Med, but they tend to be focused on single areas of specialization (biomedical research in the case of PubMed) or run by publishers. Google Scholar will probably avoid those issues.

Now, if you'll excuse me, I'm just going to wipe the drool off the corners of my mouth now. OoooOOoooooh, it even links to articles that cite the ones you're interested in....