Showing posts with label profession. Show all posts
Showing posts with label profession. Show all posts

Monday, September 29, 2008

For every thing, there is a season

Every time a colleague asks the question, "What the hell happened to computer science enrollment?" I answer, "These things are cyclical. CS enrollment will come back, and some other field's enrollment will drop." And, indeed, as the dot-com bust has faded in new student memories, enrollment has increased again (new enrollment was up last fall in my department and is up again (over 30%, year over year) this fall.

The title above links to a Computerworld article speculating that the current banking crisis will mean further increases in CS enrollment. That many or may not be true; I'd prefer the steady growth associated with students following their interests than the sudden boom of students chasing the latest fad. The article does at least have some entertainment value, with Carly Fiorina (of HP and now McCain/Palin campaign fame) showing her lack of insight that American workers are also HP (or any company's) customers -- no jobs for Americans mean no customers.

Then again, I understand this, but I'm not getting $42 million for any of my failed efforts.

Wednesday, February 27, 2008

Needing "school math" without using it

I was thinking about my previous posts about the UW College of Education's (CoE's) recent political polemic about so-called reform math. One of their major points is that engineers don't use what they call "school math": they just use computers. Please allow me to outline my own work, which is highly compute intensive and rarely involves what they would call school math, but which nevertheless I could not do without a healthy dose of school math -- not just in my education, but also in my work.

My research is in the area of computational neuroscience, in which I build mathematical models of individual nerve cells (neuron)s and groups of neurons, develop simulation software for single computers and clusters of computers, and analyze data from simulations and also from experiments on actual living tissue. This sort of work is very much like that done by anyone simulating physical systems, be they biological, chemical, mechanical, electrical, etc.

Like the engineers described in the CoE's publication, my work is heavily computational, as it isn't feasible to do this work with pencil and paper, as in "school math". The basics of the mathematical models involve a number of differential equations: equations that describe how some part of the system changes in response to other parts of the system. Now, it turns out that differential equations is covered by a pretty much standard college sophomore mathematics course. So, why isn't the stuff I do "school math"? It's simple:

  1. We only cover the mast basic type of differential equations in that class, linear equations. These are actually quite good for describing simple systems: electrical circuits made up of resistors and capacitors, mechanical systems with springs, etc. The advantage of these equations is that we can solve them on paper and they're easy to learn. The disadvantage is that they aren't very good descriptions of complicated systems like neurons (and many other, nonlinear systems). Once we move to nonlinear systems, we almost certainly need to use computers to do numerical simulations.
  2. We mostly just solve single equations in that class (there are other classes where we learn to solve groups of differential equations, later on in the curriculum). The systems I'm interested in can have hundreds or thousands of differential equations, and so I have no choice but use computer simulation.

If you were to watch me work, you would see the following (between the long periods of time in meetings, in class, preparing for said things): I decide on a question I'd like to answer, such as how the behavior of a network of neurons changes as some parameter (think: "tuning knob") is changed. I set up the parameters for a simulation or maybe bunch of simulations and, anytime from a few minutes to a few days later, I have some results. I load those results into MATLAB (numerical mathematics software) and plot the results. I then either exclaim, "Wow!", and hurriedly start writing a summary and thinking of what else I need to do to finish telling the story for a publication (rare), or I say, "Nuts!" and think again why the system either displays uninteresting behavior (Who knows; maybe its lack of interest is in itself something noteworthy? Or is that just wishful thinking?) or doesn't behave like the living nervous system. So, the observer sees that I don't "do" "school math". End of story?

Well, not quite. Because the observer doesn't see what's going on "behind the scenes" (i.e., in my mind). First of all, I would have no hope of even being able to start understanding what simulations I need to run without a very firm and extensive "school math" background. For instance, I work with a number of bright undergraduate students in my research. Some of them have math backgrounds that include differential equations and beyond and some don't. This has nothing to do with how smart they are; math beyond calculus isn't required for computer science and so only those students who come to us via a "nonstandard" pathway (e.g., changed major, previous degree/career) will have the more advanced math. Though all of these students can help out in my research, only those with more advanced "school math" are able to understand the underlying mathematical model well enough to mess with that aspect of the project (unless I teach a student some of the required "school math"). After all, unless you want to resort to randomly poking something just to see what it does, you really have to understand what's going on inside it; that's the only way you can intelligently select what kind of "poking" is likely to tell you something interesting. In fact, it's the only way you can begin to ask questions about the system, let alone start formulating experiments to answer those questions.

Even after the simulations are over, I still need to interpret the results, and this requires yet more "school math" running around in my head. What kind of result did I get? What relationship does it have to previous results I've gotten, or for that matter, results others may have gotten? What does this result mean in the overall context of the system in question and the thinks I'd like to know about it/do with it? And so on.

In other words, it is emphatically not the case that the computer has relieved me of the need to know math. All the computer has done is take over the grunt work: it has become an additional tool in my mathematical arsenal. But the computer can't think, and that thinking is where all the "school math" is. It just isn't apparent to the observer because I know it well enough that it happens in my brain automatically. This is no different than the automaticity with grammar that we use in everyday life. Just because we don't carefully label each of our utterances with "subject", "verb", "object" doesn't mean that grammar isn't necessary.

Finally, does this apply to "everyday" engineers, or just people doing research? Of course it applies to engineers (at least those who haven't "moved up" to management)! That's why businesses hire engineers: they need people who can think about solutions to problems and have the depth of background to understand the interrelationships among parts of solutions from "first principles" on up to final product. Some tasks may become routine and thus almost automatic or thoughtless, but its important to have someone who can look at a problem (or a solution proposed by some software) and say, "Wait a minute; something's fishy here." And, in the final analysis, that's the most important contribution of "school math": it is the language of creativity.

Wednesday, October 10, 2007

Measuring science

This post was inspired by an excellent one by GrrlScientist, linked from the title above. She starts off discussing journal impact factors, which are a measure of the average number of times a paper in a journal is cited by others. Then there's what is essentially a personal impact factor, which is the number of times a particular researcher's papers are cited. These have problems, which the H-index is meant to address. Briefly, a person has an H-index of h if he or she has at least h papers cited at least h times. So, if I have 100 papers, each cited once, then I have an H-index of 1. If 99 are cited once and one is cited 43,000 times, my H-index is still 1. If 95 are cited once and the remaining 5 are each cited at least 5 times, then I have an H-index of 5. And so on.

So, first of all, there is the question of gaming the system. It's unlikely that I can convince 43,000 of my colleagues to cite one of my papers (but, if you'd like, pick one from my CV on my UW home page and cite away). But if I'm only shooting for, say, an H-factor of 20 or so, then that might be doable. Supposedly, people do try to game the system by doing things like citation swapping, though this seems to me to represent time better spend being a more productive researcher (rather than just trying to look more productive or impactful).

Though I may be unconvinced about the effects of such gaming, I see this as a fatal flaw of any attempt to extract a simple metric from the interrelationships among such publications. Just look at how much effort Google has expended on providing good search results. Since these results are presented in a sequence, presumably from most relevant (or "best") to least, they have been implicitly assigned a single measure. And there's a cottage industry surrounding pushing sites' ratings up that has nothing to do with their content. I'll come back to this idea of creating a one-dimensional ordering later.

To me, there's another problem with metrics such as this. Let's say that my H-index is 11, as computed using Google Scholar. Furthermore, let's assume that issues such as self-cites (citing one's own work) and co-cites (citing of one's work by collaborators; I'll revisit this topic, too) don't effect rankings (these may be invalid assumptions). There's still one problem: is an H-index of 11 good? Bad? Middling? If we read Wikipedia, we learn, "In physics, a moderately productive scientist should have an h equal to the number of years of service while biomedical scientists tend to have higher values."

But what about computer scientists? We could consult a listing like the CS Meta H index. We would then have to compare my H-index with other faculty at similar stages in their careers who are working at similar institutions and who have had roughly similar career paths. Unfortunately, that information isn't in the index. We need to know a lot about different universities, different CS departments, and individual faculty. Maybe it would just be easier to read one or two of my papers and judge for yourself.

Coming back to the subject of co-cites, this could be considered a sign of an attempt to game the system. On the other hand, it would make more sense for me to make gaming arrangements with colleagues with whom I have no direct professional connection. (Hmm. Three more strategically placed citations will get me to an H-index of 12; five more in just the right spots will get me to 13.) But what about people who collaborate widely? Their papers will have lots of co-cites, but their work will also be more broadly influential because of all that collaboration. So, when I've prepared materials for external review, I always separate out the co-cites. Make of them what you will.

The desire to create this scalar (one-dimensional) metric of scholarly is a natural one. When I look at the complex dynamical behavior of a neural network, one of the first things I want to do is extract a single measure to characterize that behavior, so that I can then more easily examine how behavior depends on various parameters. But I have a very carefully defined question in mind when I do this. When we measure science, what is our question? Are we asking if a particular scientist is "good"? What is good? Does it mean that the scientist's work has impact in the field? How can we really ascertain this without understanding the field and the scientist's contributions in that context?

Einstein had four papers that changed the field of physics forever. But that's just an H-index of 4. I was discussing this with one of my colleagues, however, and his opinion was that 4 was a reasonable assessment of Einstein, and that we should want to hire and promote scientists who are consistently productive, not ones who have one brilliant flash of insight and then nothing approaching that for the rest of their lives. But how can we tell the difference between consistent, quality productivity and a laser-like focus on getting out each least publishable unit? To me, the only solution is knowing the person; we can't reduce the behavior of that large a neural network to a single useful measure.

Thursday, April 26, 2007

Cause and effect

Some articles in the latest round of debate over the future of the computing profession:

Which is cause and which is effect? Decreasing numbers of students interested in computing? Unpleasant working conditions, compared to other professions, many of which having less onerous coursework? Increasing immigration and outsourcing preventing salaries from increasing? Likely, this involves at least one feedback loop; I am concerned that the feedback will make matters worse, not better.

Wednesday, April 18, 2007

Women and the computing profession

The title above links to yet another article, this time from The New York Times, on the decreasing number of women entering the profession. If you've been reading articles like this, you'll likely detect the slow evolution of the message to include more assertions that demand for graduates is much higher than is perceived. All of the hard data I've seen supports the assertion that demand for computer professionals is higher now than at the peak of the dot-com boom.

One curious item, on page two, is the note that the University of Washington, Seattle "never had a programming requirement." Perhaps what was meant was for freshman admission? Because the introductory CS sequence, 142 and 143 (and the equivalent courses for transfer students) are pretty typical intro to programming classes.

Update: Here's a link the UW Seattle CSE department's "Why Choose CSE?" web site mentioned in the article.

Wednesday, March 07, 2007

Bill Gates eats his cake

Bill Gates was in Washington today to testify in front of the Senate Committee on Health, Education, Labor, and Pensions. In his testimony (PDF), he said:

A top priority must be to reverse our dismal high school graduation rates – with a target of doubling the number of young people who graduate from high school ready for college, career, and life – and to place a major emphasis on encouraging careers in math and science.
He also said:
College and graduate students are simply not obtaining science, technology, engineering, and mathematics (“STEM”) degrees in sufficient numbers to meet demand. The number of undergraduate engineering degrees awarded in the United States fell by about 17 percent between 1985 and 2004.

Unfortunately, he then goes on to say:

...the terrible shortfall in our visa supply for the highly skilled stems not from security concerns, but from visa policies that have not been updated in over a decade and a half... I personally witness the ill effects of these policies on an almost daily basis at Microsoft. Under the current system, the number of H1-B visas available runs out faster and faster each year... Barring high-skilled immigrants from entry to the U.S., and forcing the ones that are here to leave because they cannot obtain a visa, ultimately forces U.S. employers to shift development work and other critical projects offshore.

So, on the one hand, Bill Gates wants more Americans to seek technology careers. On the other hand, he wants the ability to hire more immigrants. From a potential student's point of view, these seem contradictory goals, the latter reducing the attractiveness of such careers by decreasing pay and security and therefore decreasing the number of students who might want to major in technology fields. From Bill's point of view, they are entirely consistent: increase the supply of labor to drive down its cost. (Boy, I hate sounding like Lou Dobbs.)

Oh, and in case you thought Bill's point of view was purely that of a disinterested philanthropist, he also lobbied for tax breaks for his and similar companies.

Topics: , .

Friday, February 23, 2007

More on the trouble with CS education

The title links to a Stanford press release in which Prof. Eric Roberts indicates that the problem underlying the current decline in CS enrollment lies with CS education. I agree with this: introductory CS classes, for instance, are probably the most in-your-face, student-unfriendly courses at a university. The article goes on, however, to say:

Universities also struggle with attracting enough computer science educators. 'In the '80s boom, there was one year in which there was one applicant for every seven open [teaching] positions, which means that six of the positions just did not get filled,' says Roberts. Today, there are more applicants than openings, but the ratio—hovering at around two to one—still stands in stark contrast to that in most humanities departments, where hundreds of applicants compete for one faculty job opening.

'I used to argue that Ph.D.s in computer science probably lowered your salary, because they opened lower paying jobs [in academia],' Roberts half jokes. 'There's an economic incentive not to teach but to go off and make your killing in the field.'

So, on the one hand, CS faculty are paid too little compared to industry. I'm happy to agree with someone who says I'm underpaid. On the other hand, there are not enough people trying to get faculty positions (two per job opening). Presumably, he wants more applicants and higher pay, not realizing that these are diametrically opposed goals. The reason pay is low is because there are more applicants than positions. If the number of applicants rose to the same level as in the humanities, then pay would fall to the level of that for humanities faculty.

Oh, well, Eric, thanks for playing anyway. We have an assortment of lovely virtual prizes for you to take home.

Saturday, November 04, 2006

The price paid for wasted cycles

The title above links to a Linux World article on the One Laptop Per Child (OLPC) project. That project aims to produce a laptop that is cheap enough to make its way into the hands of schoolchildren around the world. the biggest obstacle is the insatiable need for computer cycles by today's software. As I've written before, a large chunk of these cycles -- if not most -- are burned to make software more profitable for software companies, not to benefit the user.

The OLPC people seem to agree with me:

Today's laptops have become obese. Two-thirds of their software is used to manage the other third, which mostly does the same functions nine different ways.
     OLPC FAQ
Fast processors and inexpensive memory have made tidy programming a low priority, Gettys [Jim Gettys, vice president, Software Engineering] said. "A lot of people in the past decade or so have gotten quite sloppy."